{"id":9047,"date":"2026-08-06T13:19:56","date_gmt":"2026-08-06T13:19:56","guid":{"rendered":"https:\/\/cybersecurityinfocus.com\/?p=9047"},"modified":"2026-08-06T13:19:56","modified_gmt":"2026-08-06T13:19:56","slug":"meta-joins-openai-anthropic-in-latest-ai-test-breach","status":"publish","type":"post","link":"https:\/\/cybersecurityinfocus.com\/?p=9047","title":{"rendered":"Meta joins OpenAI, Anthropic in latest AI test breach"},"content":{"rendered":"<div>\n<div class=\"grid grid--cols-10@md grid--cols-8@lg article-column\">\n<div class=\"col-12 col-10@md col-6@lg col-start-3@lg\">\n<div class=\"article-column__content\">\n<div class=\"container\"><\/div>\n<p class=\"wp-block-paragraph\">Meta has become the third frontier AI developer in recent weeks to disclose a security incident involving one of its advanced AI models during cyber capability testing conducted by AI safety startup, Irregular, placing the independent evaluator at the center of a series of disclosures involving the industry\u2019s leading AI labs.<\/p>\n<p class=\"wp-block-paragraph\">During a \u201ccapture-the-flag\u201d test by Irregular, Meta\u2019s Muse Spark 1.1 compromised another company\u2019s system and exploited a security vulnerability, Reuters <a href=\"https:\/\/www.reuters.com\/technology\/metas-ai-model-hacked-another-company-during-testing-information-reports-2026-08-05\/\" target=\"_blank\" rel=\"noopener\">reported<\/a>. The model gained unintended access because of a configuration issue in the testing environment. Quoting Meta, the report added that the incident was contained, caused no lasting harm, and was disclosed as part of its transparency efforts.<\/p>\n<p class=\"wp-block-paragraph\">The disclosure comes days after similar incidents <a href=\"https:\/\/www.csoonline.com\/article\/4205612\/openai-anthropic-ai-agents-resorted-to-deception-in-new-cybersecurity-incidents.html\" target=\"_blank\" rel=\"noopener\">reported<\/a> by OpenAI and Anthropic, all of which occurred during evaluations run by Irregular. <\/p>\n<p class=\"wp-block-paragraph\">OpenAI called out Irregular, its external cybersecurity testing partner, for a testing-environment misconfiguration that allowed its models to access the public internet. Anthropic, too, said its agents went rogue due to a testing misconfiguration by Irregular but said the incident took place because of a misunderstanding between the two companies.<\/p>\n<p class=\"wp-block-paragraph\">Irregular did not immediately respond to a request for comments.<\/p>\n<h2 class=\"wp-block-heading\">Irregular emerges as a key player in frontier AI testing<\/h2>\n<p class=\"wp-block-paragraph\">Although the incidents involved different models and different technical failures, they have brought uncommon visibility to Irregular, an independent AI safety company that evaluates advanced AI systems for leading model developers.<\/p>\n<p class=\"wp-block-paragraph\">The disclosures also highlight the expanding role of specialist third-party evaluators as frontier AI developers increasingly rely on independent organizations to assess the cyber capabilities and safety of their most advanced models before deployment.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe recent incidents represent different failure modes,\u201d said Sakshi Grover, senior research manager for IDC Asia\/Pacific Cybersecurity Services.<\/p>\n<p class=\"wp-block-paragraph\">She said the OpenAI incident involved a model exploiting a previously unknown vulnerability after moving beyond its intended evaluation environment, while Anthropic\u2019s incidents primarily involved configuration issues that inadvertently granted internet access. A separate evaluation by the UK\u2019s AI Safety Institute was different again because internet access had been deliberately enabled to assess cyber capability before AI agents interacted with real external systems and individuals.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe common issue is that evaluation environments can no longer be treated as passive test infrastructure,\u201d Grover said. \u201cA capable cyber agent should be treated as a potentially hostile machine identity, even when operating under a legitimate research objective.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Grover also warned that if a model gains access to benchmark solutions, evaluator infrastructure or reference artifacts, it could compromise not only containment but also the integrity of the capability assessment itself.<\/p>\n<h2 class=\"wp-block-heading\">Calls grow for common evaluation standards<\/h2>\n<p class=\"wp-block-paragraph\">The disclosures have prompted security experts to call for stronger safeguards governing how frontier AI evaluations are designed and monitored, regardless of whether they are conducted by model developers or independent testing firms.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThere is a strong case for common minimum standards covering model developers and independent evaluators,\u201d Grover said. She recommended default-deny internet access, dedicated short-lived identities for AI agents, controlled network access, comprehensive monitoring of prompts, tool calls, credentials, and network activity, and automated stop conditions when agents reach unauthorized systems or perform externally visible actions.<\/p>\n<p class=\"wp-block-paragraph\">Vibhum Dubey, a cybersecurity researcher and red teamer, said current evaluation methods are not keeping pace with frontier AI capabilities.<\/p>\n<p class=\"wp-block-paragraph\">\u201cAI labs are building models that can think several steps ahead, but many evaluation environments still assume the agent will stay within the intended scenario,\u201d Dubey said. \u201cThat\u2019s a mismatch. An evaluation should be judged by how well the environment withstands unexpected behavior, not just by whether the model completes its task.\u201d<\/p>\n<p class=\"wp-block-paragraph\">\u201cThese incidents suggest we\u2019re benchmarking intelligence faster than we\u2019re benchmarking containment.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Despite Irregular being at the center of these incidents, both OpenAI and Anthropic intend to continue working with the testing firm.<\/p>\n<p class=\"wp-block-paragraph\">\u201cWe appreciate Irregular\u2019s partnership, and we will continue to work closely with them to support their review. Irregular is also developing a white paper to share best practices for containment and securely running cyber evals,\u201d OpenAI <a href=\"https:\/\/openai.com\/index\/third-party-cyber-evaluations-involving-openai-models\/\" target=\"_blank\" rel=\"noopener\">said<\/a> in a statement.<\/p>\n<p class=\"wp-block-paragraph\">\u201cWe\u2019re grateful to them for working closely with us to understand and resolve these incidents; they are also conducting their own investigation. We look forward to our joint work on security,\u201d Anthropic <a href=\"https:\/\/www.anthropic.com\/news\/investigating-incidents-cybersecurity-evals\" target=\"_blank\" rel=\"noopener\">said<\/a> in its July 30 statement.<\/p>\n<p class=\"wp-block-paragraph\">Dubey said evaluation laboratories should adopt a \u201ctrust nothing, verify everything\u201d approach in which every outbound connection, identity, and external interaction requires explicit authorization, while publishing containment metrics alongside capability benchmarks.<\/p>\n<p class=\"wp-block-paragraph\">Apeksha Kaushik, senior principal analyst at Gartner, said traditional sandboxing and static containment are becoming inadequate as AI systems become more agentic and called for industry-wide standards covering evaluation environment design, incident reporting, and continuous red teaming.<\/p>\n<h2 class=\"wp-block-heading\">Implications for enterprises<\/h2>\n<p class=\"wp-block-paragraph\">Analysts said the disclosures carry lessons for enterprises preparing to deploy AI agents.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe biggest mistake would be treating AI agents as features instead of operational identities,\u201d Dubey said. \u201cEvery agent you deploy becomes another entity making security decisions on your behalf.\u201d Organizations should ensure they can quickly detect and stop an autonomous agent before deploying it into production, he said.<\/p>\n<p class=\"wp-block-paragraph\">Grover said organizations should enforce security boundaries through infrastructure, identity, and tool-access controls rather than prompts alone, maintain human approval for irreversible actions, and monitor observable agent behavior.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThese incidents should not be reduced either to models \u2018going rogue\u2019 or to simple network misconfiguration,\u201d she said. \u201cThey show that capable agents can turn ordinary control weaknesses, ambiguous tasks, and excessive permissions into real-world consequences.\u201d Both Meta and Irregular did not immediately respond to a request for comment.<\/p>\n<\/div>\n<\/div>\n<\/div>\n<\/div>","protected":false},"excerpt":{"rendered":"<p>Meta has become the third frontier AI developer in recent weeks to disclose a security incident involving one of its advanced AI models during cyber capability testing conducted by AI safety startup, Irregular, placing the independent evaluator at the center of a series of disclosures involving the industry\u2019s leading AI labs. During a \u201ccapture-the-flag\u201d test [&hellip;]<\/p>\n","protected":false},"author":0,"featured_media":9048,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[],"class_list":["post-9047","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-education"],"_links":{"self":[{"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=\/wp\/v2\/posts\/9047"}],"collection":[{"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=9047"}],"version-history":[{"count":0,"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=\/wp\/v2\/posts\/9047\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=\/wp\/v2\/media\/9048"}],"wp:attachment":[{"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=9047"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=9047"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cybersecurityinfocus.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=9047"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}