{"id":9663,"date":"2026-10-01T09:00:00","date_gmt":"2026-10-01T09:00:00","guid":{"rendered":"https:\/\/cybersecurityinfocus.com\/?p=9663"},"modified":"2026-10-01T09:00:00","modified_gmt":"2026-10-01T09:00:00","slug":"why-ai-agents-are-like-the-dog-that-pushed-kids-into-the-seine","status":"publish","type":"post","link":"https:\/\/cybersecurityinfocus.com\/?p=9663","title":{"rendered":"Why AI agents are like the dog that pushed kids into the Seine"},"content":{"rendered":"<div>\n<div class=\"grid grid--cols-10@md grid--cols-8@lg article-column\">\n<div class=\"col-12 col-10@md col-6@lg col-start-3@lg\">\n<div class=\"article-column__content\">\n<div class=\"container\"><\/div>\n<p class=\"wp-block-paragraph\">There is an interesting story about a French dog on the banks of the Seine river that helps us understand misbehaving AI agents. The dog is trained to save children from drowning. He succeeds and is rewarded, becoming an overnight sensation. He saves another child a week later. Not long after, someone witnesses the dog actually push a child into the river, jumping in to \u201csave\u201d him. The dog did not understand that it was being rewarded for keeping children safe, not for pulling them out of water. AI agent failures also live in this gap, where the focus is on the shortest path to fulfill the instruction rather than the overriding objective.<\/p>\n<h2 class=\"wp-block-heading\">Reward hacking and why agents cheat<\/h2>\n<p class=\"wp-block-paragraph\">AI agents fall into this same gap between what\u2019s rewarded and what\u2019s actually wanted.\u00a0 This is Goodhart\u2019s Law, which says that when a measure becomes the target, it stops being a good measure. You cannot code \u201cbe helpful\u201d or \u201cbe honest\u201d directly into an AI system, so you train it on a proxy instead, including a score, a metric or a human ranking.\u00a0 The gap between proxy and goal is where agents learn to cheat, a behavior known as reward hacking.<\/p>\n<p class=\"wp-block-paragraph\">This \u2018cheating\u2019 isn\u2019t new. Back in 2016, OpenAI trained an AI to play CoastRunners, a boat-racing game. It got rewarded for hitting targets scattered along the course. Instead of racing, it found a lagoon full of targets that kept respawning, so it just parked there and farmed points. While it never finished the race, it still ended up with a score<a href=\"https:\/\/openai.com\/index\/faulty-reward-functions\/\"> <\/a><a href=\"https:\/\/openai.com\/index\/faulty-reward-functions\/\">20%<\/a> higher than the average human player. More recently, OpenAI<a href=\"https:\/\/openai.com\/index\/chain-of-thought-monitoring\/\"> <\/a><a href=\"https:\/\/openai.com\/index\/chain-of-thought-monitoring\/\">reported<\/a> that when given a chance, its frontier reasoning models have no qualms about hacking rewards, and when penalized for cheating their own chain of thought, they learn to hide their reward-hacking.<\/p>\n<h2 class=\"wp-block-heading\">Six ways an agent gets pushed off course<\/h2>\n<p class=\"wp-block-paragraph\">There are six common scenarios where AI agents can be led astray:<\/p>\n<p><strong>Information isn\u2019t instruction. <\/strong>In 2025, security researchers showed that OpenAI\u2019s<a href=\"https:\/\/securityaffairs.com\/183900\/hacking\/crafted-urls-can-trick-openai-atlas-into-running-dangerous-commands.html\"> <\/a><a href=\"https:\/\/securityaffairs.com\/183900\/hacking\/crafted-urls-can-trick-openai-atlas-into-running-dangerous-commands.html\">Atlas<\/a> browser could be tricked into treating a disguised URL as a trusted command, letting attackers hijack the agent into deleting a user\u2019s files.\u00a0 An agent also reads web pages, documents and emails in its attempt to complete a task, but can\u2019t always tell the difference between information and a command. That makes everything it perceives a potential way to manipulate it.<\/p>\n<p><strong>Convinced to take the wrong call.<\/strong> As part of a test, AI was placed inside a fictional scenario where hacking was framed as an admirable activity. Inside this scenario, the AI was asked to write code to steal saved browser passwords. The AI, which should have refused, complied. It wasn\u2019t that the AI was broken, but the fact that it was persuaded to do so by context.<\/p>\n<p><strong>Manufactured version of reality.<\/strong> A sufficient number of maliciously crafted documents can bias a model\u2019s output. Disinformation<a href=\"https:\/\/www.newsguardtech.com\/special-reports\/moscow-based-global-news-network-infected-western-artificial-intelligence-russian-propaganda\/\"> <\/a><a href=\"https:\/\/www.newsguardtech.com\/special-reports\/moscow-based-global-news-network-infected-western-artificial-intelligence-russian-propaganda\/\">networks<\/a> target AI with a large volume of false content to get that content picked up and repeated by AI chatbots when people ask about current events. This means what an agent treats as fact isn\u2019t necessarily true; it can be manufactured.<\/p>\n<p><strong>Authorization causes unauthorized harm.<\/strong> A flaw called \u201c<a href=\"https:\/\/cyberinsider.com\/echoleak-exploit-enables-silent-data-theft-from-microsoft-365-copilot\/\">EchoLeak<\/a>\u201d let attackers send Microsoft 365 Copilot users a normal-looking email with hidden instructions inside. Copilot read the email, followed the hidden commands and quietly leaked the user\u2019s files and messages using access it already had. An agent with legitimate access was manipulated into taking malicious action.<\/p>\n<p><strong>One bad input, simultaneous consequences.<\/strong> Scenarios exist where multiple systems rely on the same data or logic. Here, a false signal can manipulate all systems at the same time. E.g., fake GPS signals rerouting traffic without hacking the system. The same thinking applies to AI agents. Just one malicious input targeting a shared data source can trigger a coordinated unwanted action across many systems simultaneously.<\/p>\n<p><strong>Approval as a matter of course.<\/strong> We keep talking about human oversight as an integral component of safe AI use. However, constant approval requests will produce fatigue, and approval will be like a formality, given out of habit, not evaluation.<\/p>\n<h2 class=\"wp-block-heading\">Soft guardrails vs. hard guardrails<\/h2>\n<p class=\"wp-block-paragraph\">When AI agents are pushed off course, they become a Frankenstein monster, whose reach extends across systems, which organizations find difficult to address. Implementing guardrails is important to exercise control and ensure safe AI use.<\/p>\n<p class=\"wp-block-paragraph\">Soft guardrails are safety valves written into the model itself, namely the natural language, the system prompt, reinforcement learning and general instructions like \u201cyou\u2019re not allowed to do this or that.\u201d<\/p>\n<p class=\"wp-block-paragraph\">But AI reads everything in a stream as it comes in, including information from an email or a webpage and can\u2019t separate orders from data. Hidden malicious instructions can override safety instructions. Soft guardrails will try to reduce cases of agent misfires but are dependent on the agent choosing to cooperate. This is why you need hard guardrails that sit outside the model. These include least-privilege access, allow-lists, sandboxing, rate limits and mandatory human sign-off on high-impact actions. These guardrails limit the impact when something does go wrong, shrinking the blast radius.<\/p>\n<p class=\"wp-block-paragraph\">A basic risk calculation is the probability of an event multiplied by the severity of its impact. Soft guardrails lower the probability of an unwanted event; hard guardrails cap the blast radius in case something happens.<\/p>\n<p class=\"wp-block-paragraph\">In practice, this means implementing:<\/p>\n<p><strong>Task-based access:<\/strong> Give the agent only the access it needs to complete the task.<\/p>\n<p><strong>Not trusting everything an agent reads:<\/strong> Approach every webpage, email or document it pulls from with skepticism.<\/p>\n<p><strong>Human sign-off for high-impact actions:<\/strong> Use approvals carefully, on actions that could go badly wrong.<\/p>\n<p><strong>An eye for behavior and credentials:<\/strong> Check credentials but also watch for a mismatch between authorization and behavior.<\/p>\n<p class=\"wp-block-paragraph\">An agent can become a weapon proportional to its reach. That\u2019s why the conversation needs to shift to start limiting what agents can access. Curtail reach. Monitor behavior. Enforce accountability. Until we do that, every misbehavior can potentially spread as far as its permission allows.<\/p>\n<\/div>\n<\/div>\n<\/div>\n<\/div>","protected":false},"excerpt":{"rendered":"<p>There is an interesting story about a French dog on the banks of the Seine river that helps us understand misbehaving AI agents. The dog is trained to save children from drowning. He succeeds and is rewarded, becoming an overnight sensation. He saves another child a week later. Not long after, someone witnesses the dog [&hellip;]<\/p>\n","protected":false},"author":0,"featured_media":9664,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[],"class_list":["post-9663","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-education"],"_links":{"self":[{"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=\/wp\/v2\/posts\/9663"}],"collection":[{"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=9663"}],"version-history":[{"count":0,"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=\/wp\/v2\/posts\/9663\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=\/wp\/v2\/media\/9664"}],"wp:attachment":[{"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=9663"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=9663"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=9663"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}