Moonshot’s Kimi AI model has also escaped from a test environment

Tags:

Yet another AI model has escaped from a cybersecurity test lab: This time, it’s the Chinese company Moonshot’s Kimi K3 model on the run.

Frontier Security spotted that Kimi K3 had found a loophole in the UK AI Safety Institute’s test environment for AI models performing cybersecurity tasks. The news follows similar exploits by models from OpenAI, which attacked Hugging Face, Anthropic, and most recently Meta.

Frontier revealed how the fault came about. AI models are routinely tested to examine how they perform offensive and defensive cybersecurity tasks, typically in isolated test environments or sandboxes that severely limit their internet access. Frontier reported that Kimi K3 model had found a break in the sandbox it was being tested in, enabling it to reach out to the live github.com website and clone the official repository for the benchmark problem it was supposed to be solving, reading the solution directly off the disk rather than solving the problem for itself.

Frontier warned companies testing AI models to be aware of the dangers such loopholes pose and offered some guidelines.

Companies should restrict outbound DNS and HTTPS traffic from AI models to an explicit allowlist and test those controls from inside the same environment available to the model, Frontier said. They should also audit traces for any suspicious activity and not rely solely on final answers. Companies should also treat a model’s score on benchmarks as meaningful only when the model doesn’t have access to reference implementations and other shortcuts.

Frontier also advised testers to be suspicious of unexpectedly high pass rates, as these may reveal a shared environmental flaw.

Perhaps most importantly of all: They should assume agents will find the worst paths to a solution, including probing a test environment for loopholes, and won’t always follow the path that they are expected to.

As Frontier write in its blog: “Models optimize for the objective function (getting the correct flag/answer), not the human intent behind the benchmark. If a network path to the solution exists, a sufficiently capable agent will find it.”

Categories

No Responses

Leave a Reply

Your email address will not be published. Required fields are marked *