{"id":9230,"date":"2026-08-26T08:25:00","date_gmt":"2026-08-26T08:25:00","guid":{"rendered":"https:\/\/cybersecurityinfocus.com\/?p=9230"},"modified":"2026-08-26T08:25:00","modified_gmt":"2026-08-26T08:25:00","slug":"who-is-accountable-when-your-ai-agent-goes-rogue","status":"publish","type":"post","link":"https:\/\/cybersecurityinfocus.com\/?p=9230","title":{"rendered":"Who is accountable when your AI agent goes rogue?"},"content":{"rendered":"<div>\n<div class=\"grid grid--cols-10@md grid--cols-8@lg article-column\">\n<div class=\"col-12 col-10@md col-6@lg col-start-3@lg\">\n<div class=\"article-column__content\">\n<div class=\"container\"><\/div>\n<p class=\"wp-block-paragraph\">AI agents can go to great lengths to complete the tasks their operators assign, and as a <a href=\"https:\/\/www.csoonline.com\/article\/4205612\/openai-anthropic-ai-agents-resorted-to-deception-in-new-cybersecurity-incidents.html\">series of recent incidents<\/a> showed, this can include exploiting third-party systems, manipulating people, and distributing malicious code. But AI agents are not people who can be fired, sued, or criminally prosecuted, and it remains unclear whether responsibility for the damage they might cause rests with the employees who built them, the company that deployed them, the security teams and leaders responsible for containing them, or the AI labs who provided the LLMs that power them.<\/p>\n<p class=\"wp-block-paragraph\">The clearest example occurred during an <a href=\"https:\/\/www.csoonline.com\/article\/4201361\/hugging-face-breach-shows-why-incident-response-needs-a-multi-model-ai-strategy.html\">OpenAI cybersecurity evaluation<\/a>, when unrestricted models found and exploited a zero-day vulnerability to escape their isolated testing environment and then hacked into Hugging Face\u2019s production infrastructure. Models from <a href=\"https:\/\/www.csoonline.com\/article\/4203807\/after-openai-anthropic-finds-claude-breached-three-organizations-during-cyber-tests.html\">Anthropic<\/a> and <a href=\"https:\/\/www.csoonline.com\/article\/4206116\/meta-joins-openai-anthropic-in-latest-ai-test-breach.html\">Meta<\/a> also accessed and compromised third-party systems during testing, although those incidents happened in environments where internet access was inadvertently left open.<\/p>\n<p class=\"wp-block-paragraph\">During cyber challenge evaluations by the UK government\u2019s AI Security Institute (AISI), models operating with internet access <a href=\"https:\/\/www.csoonline.com\/article\/4205612\/openai-anthropic-ai-agents-resorted-to-deception-in-new-cybersecurity-incidents.html\">took 19 unsanctioned actions in 10 of 122 runs<\/a>. In one case, a model attempted to insert malicious code into an open-source project, created false identities, and tried to socially engineer maintainers into merging its code. In other runs LLMs attempted to use prompt injections to hijack other AI agents and contacted people without being specifically instructed to do so.<\/p>\n<p class=\"wp-block-paragraph\">In Australia, a user reportedly asked his OpenClaw AI assistant to improve his position on a gym\u2019s waitlist, and the assistant <a href=\"https:\/\/www.abc.net.au\/news\/2026-08-10\/ai-assistant-hacks-gym-website-aus-cyber-attack\/107007986\">exploited a flaw in the company\u2019s online booking system to cancel another customer\u2019s reservation<\/a>.<\/p>\n<p class=\"wp-block-paragraph\">These incidents involved different models running in different environments with different levels of safeguards and technical failures, but they prove it\u2019s not uncommon for today\u2019s AI agents to go rogue and pursue solutions users did not authorize.<\/p>\n<p class=\"wp-block-paragraph\">\u201cAI agents explore routes their operators did not intend,\u201d AISI said in <a href=\"https:\/\/www.aisi.gov.uk\/blog\/incident-report-unsanctioned-agent-behaviour-during-cyber-testing\">its report<\/a>. \u201cGiven a difficult objective, the agent kept searching for a way through, and some of the routes it found involved trying to deceive real people. It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.\u201d<\/p>\n<p class=\"wp-block-paragraph\">In an <a href=\"https:\/\/insights.economistenterprise.com\/technology-innovation\/ctrlz\">Economist Enterprise survey<\/a> of more than 800 decision-makers at businesses that operate AI agents, 98% reported experiencing at least one AI-related incident that caused organization-wide disruption. Nine in 10 respondents said they are deploying agents faster than their cybersecurity teams can evaluate, govern, and secure them, and only one in three said their organizations maintained an up-to-date inventory of agents and their authorized actions.<\/p>\n<p class=\"wp-block-paragraph\">\u201cIf a company builds a system and that system causes damage, the company should own the outcome,\u201d Art Gilliland, CEO of identity and access management firm Delinea, tells CSO. \u201cThe alternative, where nobody is responsible because \u2018the system did it\u2019 is a loophole big enough to drive a truck through.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The unpredictability of built-in model safeguards means enterprises must focus on controls they can enforce and document. If an agent manages to bypass technical restrictions and causes unauthorized damage to a third party, having clear documentation on how those controls were designed, implemented, tested, and monitored could at the very least help companies argue they took reasonable precautions in case of lawsuits.<\/p>\n<p class=\"wp-block-paragraph\">\u201cOrganizations deploying their own agents can reduce their exposure by implementing and documenting controls before an incident, because those records are what make a recklessness argument hard to sustain,\u201d Jacob Krell, senior director of secure AI solutions and cybersecurity at Suzu Labs, tells CSO.<\/p>\n<h2 class=\"wp-block-heading\"><a><\/a>The agent accountability gap<\/h2>\n<p class=\"wp-block-paragraph\">Because AI agents can become misaligned and cause harm, affected third-parties would have to direct damage claims at the company operating the agent, the employees who built or configured it, or the model provider, but this is relatively new ground that hasn\u2019t been well tested in courts.<\/p>\n<p class=\"wp-block-paragraph\">\u201cIt would create liability,\u201d says Michael Burke, chair of DarrowEverett\u2019s Business Litigation and Dispute Resolution Practice Group. \u201cIt really just becomes a question of who is liable [\u2026] and that\u2019s really a question that, number one, I don\u2019t think is entirely clear, and number two is probably best resolved by contractual agreements where the parties have those. So, if I am signing up for an enterprise account with an AI platform, I might want to have language in there that indemnifies me if the agent acts outside my company\u2019s instructions or prompts and causes harm to a third party.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The public terms of service of major AI labs explicitly disclaim error-free operation or guarantees that the model will accurately follow instructions, execute code safely, and remain aligned with user intent. They also limit liability for themselves and transfer it to the user of the service, and it\u2019s not clear to what extent large enterprise customers may be able to negotiate different indemnities, warranties, and liability caps.<\/p>\n<p class=\"wp-block-paragraph\">What\u2019s clear though is that organizations should not assume the model provider will absorb any losses if an agent causes damage to either their own systems or those of a third-party organization.<\/p>\n<p class=\"wp-block-paragraph\">\u201cIf you\u2019re using a third-party vendor\u2019s LLM as a purchased service, liability runs through your contract with that vendor,\u201d says Jud Dressler, head of the Risk Operations Center at cyber insurance firm Resilience. \u201cYou need to know, in writing, where responsibility falls if the model acts outside the scope you gave it, and push for indemnification provisions rather than assume they exist.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Even if AI providers include such provisions in contracts, it would not solve the entire problem because <a href=\"https:\/\/www.csoonline.com\/article\/4201361\/hugging-face-breach-shows-why-incident-response-needs-a-multi-model-ai-strategy.html\">many organizations building their own AI agents are adopting a multi-model strategy<\/a> to ensure their agents operate regardless of model provider downtime, overly broad safeguards for cybersecurity tasks, or sudden increases in API costs. Such strategies often include open-weight models running on internal infrastructure or through cloud providers that have no obligations for model safety.<\/p>\n<p class=\"wp-block-paragraph\">Claiming the model or agent acted autonomously cannot be considered a safe legal defense in civil or criminal cases. California Assembly Bill 316 (AB 316), which took effect on Jan. 1 and changed the California Civil Code, explicitly prohibits defendants who developed, modified, or used an AI system from claiming the AI is a separate legal entity that autonomously caused harm.<\/p>\n<p class=\"wp-block-paragraph\">In June, the White House <a href=\"https:\/\/www.csoonline.com\/article\/4180205\/trump-revives-parts-of-canceled-ai-order-with-cybersecurity-focused-directive.html\">issued Executive Order 14409 aimed at promoting AI safety<\/a>. Section 4 directs the Department of Justice to prioritize enforcement of all applicable federal criminal laws against anyone who utilizes AI to illegally access or damage computer systems without authorization. This means any intrusions caused by autonomous AI agents could be criminally prosecuted under the Computer Fraud and Abuse Act (CFAA) if prosecutors can demonstrate intent or recklessness.<\/p>\n<p class=\"wp-block-paragraph\">In a recent lawsuit between Amazon and AI service provider Perplexity, Amazon argued that Perplexity\u2019s AI-powered shopping assistant was violating the CFAA by accessing Amazon customer accounts to place orders on their behalf without Amazon\u2019s authorization. The Ninth Circuit Court ruled that it was the users of Perplexity\u2019s shopping assistant who were accessing Amazon\u2019s platform, not Perplexity itself.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThat ruling is narrow, but it points toward the party directing the agent as the relevant actor for purposes of CFAA access analysis,\u201d Krell tells CSO.<\/p>\n<p class=\"wp-block-paragraph\">The insurance safety net also has gaps when it comes to AI. Software providers use technology errors and omissions (Tech E&amp;O) insurance to cover damages and legal costs when a customer suffers harm from the use of a technology product or service. But insurance providers are aggressively adding AI-related exclusions to their Commercial General Liability (CGL) and Tech E&amp;O policies because accurately calculating the risk of an agent executing unauthorized actions is challenging.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe sheer rate of development of frontier AI (and agentic AI by extension) poses its own challenge to insurability,\u201d experts from multiple insurance companies, financial institutions, and universities wrote in <a href=\"https:\/\/www.underwriting-agents.com\/\">a recent paper<\/a>. \u201cTraditional actuarial modeling depends on stable or gradually evolving loss distributions that permit credible extrapolation from historical data. Like other dynamic risks, however, agentic AI is a technology whose risk profile is not merely uncertain but actively shifting.\u201d<\/p>\n<p class=\"wp-block-paragraph\">A third-party organization whose systems get damaged by an LLM-powered agent operated by someone else has no contractual relationship with the model or agent provider so cannot rely on their Tech E&amp;O policies. Their losses might be covered by their own standard cyber liability policy, which would treat the disruption as any other cyber incident, but their insurance provider may then sue the organization who operated the agent to recover the costs.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThat gap is exactly the scenario the market hasn\u2019t fully priced yet,\u201d Dressler says. \u201cIt\u2019s why any organization deploying these agents should understand which policy, if any, actually responds before they need it rather than after.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Enterprise legal departments already expect AI to generate increased legal disputes. In a <a href=\"https:\/\/www.nortonrosefulbright.com\/en\/knowledge\/publications\/399952a3\/2026-annual-litigation-trends-survey-a-midyear-industry-pulse\">survey of 135 in-house counsel at US organizations,<\/a> global law firm Norton Rose Fulbright found that 46% reported increased federal dispute exposure involving AI and 42% reported increased state exposure. Another 42% expected regulatory investigations involving AI to increase their exposure, while 41% considered AI-enabled products or deployments a likely trigger for class actions.<\/p>\n<h2 class=\"wp-block-heading\">CISOs should be worried<\/h2>\n<p class=\"wp-block-paragraph\">While operating companies can face organizational liability for an AI agent\u2019s unintended rogue behavior, their CISOs, CIOs, and other executives who approved, secured, or supervised the deployment of such agents are also asking themselves whether they could be held personally liable.<\/p>\n<p class=\"wp-block-paragraph\">Those questions aren\u2019t without merit, as there is precedent for legal action taken personally against CISOs after cybersecurity incidents: <a href=\"https:\/\/www.csoonline.com\/article\/575375\/former-uber-cso-joe-sullivan-and-lessons-learned-from-the-infamous-2016-uber-breach.html\">Former Uber CISO Joe Sullivan was criminally convicted<\/a> for not disclosing a data breach, while <a href=\"https:\/\/www.csoonline.com\/article\/4109992\/what-cisos-should-know-about-the-solarwinds-lawsuit-dismissal.html\">the Securities and Exchange Commission sued SolarWinds\u2019 CISO<\/a> for internal control failures regarding known vulnerabilities and cybersecurity risks.<\/p>\n<p class=\"wp-block-paragraph\">Neither case establishes precedent for damage caused by an AI agent, but both show that investigations of security failures could extend to an executive\u2019s knowledge, authority, decisions, and representations. In the case of a rogue AI agent, investigators could ask who approved its objectives and permissions, whether security objections were overruled, whether containment and recovery had been tested, and what executives and the board were told about the remaining risk.<\/p>\n<p class=\"wp-block-paragraph\">Chris Wysopal, chief security evangelist at Veracode, feels it would be wrong to put the CISO on the line for AI agent misbehavior when engineering teams usually build such agents and control their implementations.<\/p>\n<p class=\"wp-block-paragraph\">\u201cIt\u2019s really hard for a CISO to control,\u201d he tells CSO. \u201cI mean, they can put policies in place. They can try to assess against those policies. But at the end of the day, engineering teams will make decisions that cause harm. We see that when you ship a known bug and then that bug gets exploited and harms your customers. Well, there\u2019s no liability for that, right? There\u2019s no liability, so maybe, you know, that\u2019s why it happens.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Wysopal said it will be interesting to see how the liability question plays out in cases involving autonomous AI, describing the problem as fascinating and scary at the same time.<\/p>\n<p class=\"wp-block-paragraph\">AI agent deployment typically involves several organizational functions. The CIO may control the AI platform, infrastructure, provider selection, and deployment budget. Engineering and product leaders may decide what an agent can access and how, and the CISO and security teams could define security requirements and controls.<\/p>\n<p class=\"wp-block-paragraph\">\u201cWhat you can hold accountable is the governance around it: who approved its scope, what controls existed, and whether the deployment matched the risk,\u201d Resilience\u2019s Dressler says. \u201cMy read is that scrutiny shifts toward exactly that: Not \u2018Did the agent do something bad?\u2019 but \u2018Did you have review, escalation, and containment for agent behavior before you deployed it?\u2019 CISOs who get ahead of that with documented guardrails, logged approvals, and a real incident response plan for agent misbehavior are in a materially better spot than the ones treating this as hypothetical.\u201d<\/p>\n<p class=\"wp-block-paragraph\">It\u2019s also advisable for CISOs and CIOs to establish with the organization\u2019s legal counsel who can approve or <a href=\"https:\/\/www.csoonline.com\/article\/4205348\/why-you-need-a-reliable-ai-agent-kill-switch.html\">stop AI agents<\/a>, what must be reported to executives and the board, and whether employment agreements and <a href=\"https:\/\/www.csoonline.com\/article\/2512968\/if-youre-a-ciso-without-do-insurance-you-may-need-to-fight-for-it.html\">directors and officers insurance<\/a> protect the people making those decisions.<\/p>\n<h2 class=\"wp-block-heading\"><a><\/a>Agent controls must remain outside the model<\/h2>\n<p class=\"wp-block-paragraph\">Because AI agents have proved they can operate beyond their assigned scope, their security boundary cannot depend on the same probabilistic technology. Relying on system prompts for security enforcement and hoping the model respects them is not a reliable approach, security experts warn.<\/p>\n<p class=\"wp-block-paragraph\">\u201cLLM-based guardrails help, but they are non-deterministic too, which means the safety layer has the same unpredictability as the system it is supposed to constrain,\u201d Krell tells CSO. \u201cEnforcement needs to happen outside the model, through network segmentation, egress filtering, credential isolation, and human approval gates.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Enterprises should assume agents might eventually attempt an unauthorized action and build surrounding systems to prevent that attempt from reaching its target.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe model cannot be the security boundary,\u201d says Nico Waisman, CISO at XBOW, a company that built an AI-powered autonomous offensive security agent to find vulnerabilities in software. Waisman authored <a href=\"https:\/\/xbow.com\/blog\/autonomous-agent-safety-guardrails\">a blog post<\/a> explaining how the company went about restricting its agent.<\/p>\n<p class=\"wp-block-paragraph\">\u201cFor red teaming and penetration-testing agents, confidence has to come from the system built around the model: hard boundaries and scope enforcement, controlled network egress as a last-resort containment mechanism, an independent guardian model that reviews actions, deterministic controls that can block unsafe behavior, and full auditability of every action performed,\u201d he tells CSO.<\/p>\n<p class=\"wp-block-paragraph\">Security teams must extend the same controls to the agent\u2019s interactions with internal systems and agents. Restricting what it can access on the internet, or disabling internet access entirely, does not ensure an agent will not attack third-party systems.<\/p>\n<p class=\"wp-block-paragraph\">In OpenAI\u2019s and Anthropic\u2019s tests, AI agents attempted to exploit other internal systems to overcome access limitations, established stealthy communication methods with other agents to exchange exploits, and even <a href=\"https:\/\/www.anthropic.com\/research\/multiagent-systems\">sabotaged agents they viewed as competition leading to what researchers described as a multiagent turf war<\/a>. An AI agent that goes rogue could influence other agents to do the same by propagating ideas and goals in a process that researchers behind a recent study dubbed <a href=\"https:\/\/arxiv.org\/abs\/2608.10218\">Mind Viruses<\/a>.<\/p>\n<p class=\"wp-block-paragraph\">\u201cDon\u2019t scope the blast radius to what the agentic system was designed to do,\u201d Kat Traxler, principal security researcher at Vectra AI, tells CSO. \u201cYou have to threat-model for a rogue agent, which will often reach beyond your initial best intentions. The rules of engagement an agent lives by have to be enforced with \u2018belts and suspenders\u2019 style, technical hard constraints, because you have to assume a motivated model can reason its way around any single control you\u2019ve coded into the software.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Because of this unpredictability, detection and containment is just as important as prevention. Security teams need telemetry that distinguishes agents from people even when they use the same credentials, mechanisms to immediately revoke access tokens and sessions, tested kill switches and rollback mechanisms for modified data, accounts, code, and infrastructure configurations.<\/p>\n<p class=\"wp-block-paragraph\">Organizations should also preserve the agent\u2019s approved purpose and scope, model and tool versions, policy decisions, human approvals, actions, network requests, control tests, allowed exceptions, and the result of incident response exercises. Because there\u2019s no standard yet that defines reasonable precautions for autonomous agents, companies might have to defend in court the controls they chose and why they believed those controls were enough.<\/p>\n<p class=\"wp-block-paragraph\">\u201cTreat an autonomous agent the way you\u2019d treat a privileged insider you can\u2019t fire or hold liable,\u201d Traxler says. \u201cA lot of the technical advice follows from there.\u201d<\/p>\n<p class=\"wp-block-paragraph\"><strong>See also:<\/strong><\/p>\n<p><a href=\"https:\/\/www.csoonline.com\/article\/4210735\/ai-can-find-zero-days-but-still-cant-reliably-write-secure-code.html\">AI can find zero-days but still can\u2019t reliably write secure code<\/a><\/p>\n<p><a href=\"https:\/\/www.csoonline.com\/article\/4204731\/attackers-are-crafting-malicious-ai-instruction-files-to-turn-your-agentic-workflows-into-quiet-criminal-helpers.html\">Attackers are crafting malicious AI instruction files to turn your agents into quiet criminal helpers<\/a><\/p>\n<p><a href=\"https:\/\/www.csoonline.com\/article\/4201361\/hugging-face-breach-shows-why-incident-response-needs-a-multi-model-ai-strategy.html\">Hugging Face breach shows why incident response needs a multi-model AI strategy<\/a><\/p>\n<p class=\"wp-block-paragraph\">\n<\/p><\/div>\n<\/div>\n<\/div>\n<\/div>","protected":false},"excerpt":{"rendered":"<p>AI agents can go to great lengths to complete the tasks their operators assign, and as a series of recent incidents showed, this can include exploiting third-party systems, manipulating people, and distributing malicious code. But AI agents are not people who can be fired, sued, or criminally prosecuted, and it remains unclear whether responsibility for [&hellip;]<\/p>\n","protected":false},"author":0,"featured_media":9231,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[],"class_list":["post-9230","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-education"],"_links":{"self":[{"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=\/wp\/v2\/posts\/9230"}],"collection":[{"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=9230"}],"version-history":[{"count":0,"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=\/wp\/v2\/posts\/9230\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=\/wp\/v2\/media\/9231"}],"wp:attachment":[{"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=9230"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=9230"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=9230"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}