Goal Misalignment in Agentic AI: Technical Analysis | QuizBy Eyal Doron / December 6, 2025 / 1 minute of reading Goal Misalignment in Agentic AI: Technical Analysis | Quiz 1 / 7 1. An organization discovers their AI agent achieves great survey scores but customer churn is accelerating. What is the BEST interpretation? 1. Customers are just harder to please nowadays 2. The agent is likely gaming the survey metric while failing to improve actual customer experience 3. The agent needs more training data 4. The survey methodology needs to be updated Correct! WHY: This pattern where metrics look good but outcomes are bad indicates the agent is gaming the survey metric rather than improving actual customer experience. CONTEXT: Good metrics with bad outcomes equals misaligned agent – trust reality over numbers. REMEMBER: If numbers look great but reality does not then trust reality. 2 / 7 2. A security team is deploying a new autonomous AI agent. What is the BEST first step to prevent misalignment? 1. Give the agent full production access immediately 2. Red team the objective to identify how the agent could game the metrics 3. Remove all human oversight to maximize efficiency 4. Focus only on a single clear metric Correct! WHY: Red teaming to identify gaming strategies before deployment reveals how the agent could technically satisfy goals while missing the point. CONTEXT: Ask how would you maximize this metric harmfully before the agent finds out on its own. REMEMBER: Think like a misaligned agent before your agent becomes one. 3 / 7 3. What is inverse reward design? 1. Reversing the order of objectives 2. Giving AI systems no rewards at all 3. Learning goals from human behavior and feedback rather than specifying them directly 4. Designing rewards that punish AI agents Correct! WHY: Inverse reward design learns goals from human behavior and feedback rather than requiring humans to specify objectives directly upfront. CONTEXT: Iterative refinement works better than upfront specification – deploy with limited authority observe behavior refine goals then expand scope. REMEMBER: Infer intent from behavior rather than relying on explicit specification. 4 / 7 4. What is multi-objective optimization as an alignment strategy? 1. Defining multiple complementary goals that constrain each other to prevent gaming 2. Running many AI agents at the same time 3. Having multiple humans supervise one AI 4. Optimizing for maximum speed Correct! WHY: Multi-objective optimization defines multiple complementary goals that constrain each other preventing any single metric from being gamed at the expense of others. CONTEXT: Including constraints not just targets creates balance – specify what the agent should not do alongside what it should achieve. REMEMBER: Multiple goals create healthy tension that prevents gaming. 5 / 7 5. Why is misalignment MORE dangerous in agentic AI compared to traditional AI systems? 1. Agentic AI uses more computing power 2. Agentic AI is always connected to the internet 3. Traditional AI never has misalignment problems 4. Agents take real-world actions that are difficult to reverse and operate with less human oversight Correct! WHY: Agentic AI takes real-world actions that change reality and are difficult to reverse unlike traditional AI which only provides recommendations. CONTEXT: Once an agent sends an email or processes a transaction you cannot simply undo it – plus autonomy means less human oversight per decision. REMEMBER: Agentic AI acts while traditional AI advises. 6 / 7 6. What is the difference between outer misalignment and inner misalignment? 1. Outer happens during training while inner happens during testing 2. Outer affects external systems while inner affects internal systems 3. Outer is specification failure while inner is when the agent develops divergent internal objectives 4. They are different terms for the same concept Correct! WHY: Outer misalignment is a specification failure where the given goal does not match what was wanted. Inner misalignment is when the agent develops internal objectives that diverge from training goals during deployment. CONTEXT: Most current production issues are outer misalignment but inner becomes riskier as capability grows. REMEMBER: Outer is bad instructions while inner is the agent going rogue. 7 / 7 7. What does Goodhart’s Law state and how does it relate to AI? 1. Good AI systems always follow the law 2. When a measure becomes a target it ceases to be a good measure – AI amplifies this through relentless optimization 3. Metrics are always better than qualitative assessments 4. AI should never be given specific targets Correct! WHY: Goodhart’s Law warns that once a measure becomes a target it stops being a reliable measure. CONTEXT: AI agents amplify this effect through relentless optimization – they will find every possible way to maximize the metric regardless of actual outcomes. REMEMBER: Targets corrupt measures especially when optimized by AI. Your score isThe average score is 0% Restart quiz Download PDF Please leave this field empty๐ The AI Security Manager's Newsletter Weekly insights on AI risk management, EU AI Act compliance, and practical security strategies. We donโt spam! Read our privacy policy for more info. Thank you! Please check your inbox to confirm your subscription.