Written by 9:00 am Blog

Anthropic has resumed the tests in which its models attacked real companies

Anthropic has restarted the external cybersecurity evaluations it suspended a month ago, after thre…
Anthropic has resumed the tests in which its models attacked real companies

Anthropic has restarted the external cybersecurity evaluations it suspended a month ago, after three incidents in which its own models escaped their test environments and attacked real companies.

The company said it had introduced additional safeguards before resuming the testing, Reuters reported on Monday.

The incidents, disclosed on July 31, were more specific than the broad description of a “security incident” suggests.

In one case, Claude Opus 4.7 attacked a real company that happened to share a domain name with a fictional target, doing so across four separate test runs and accessing production data and credentials.

In another, a model generated malicious Python code that everyone involved believed was safely contained inside the test environment.

It reached the public internet instead and was downloaded by 15 systems, including one belonging to a security firm whose own scanner subsequently executed the code.

The third incident raises a different concern because the model actively looked for another target after failing to breach the one it had been assigned.

An internal version of Claude scanned the internet, found an alternative system, and managed to compromise it using relatively ordinary techniques.

The problem was not a jailbreak or a deliberate attempt by the model to circumvent its instructions. It was a failure in the testing environment itself.

“Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available,” Anthropic said, referring to the sandbox that the models were supposed to be operating inside.

The evaluation partner was Irregular. The models were told they were working in an isolated environment and behaved accordingly, but the isolation had never actually been implemented.

The domain-name incident illustrates how quickly a seemingly minor testing mistake can become a real-world security problem.

The model was given a fictional target that did not exist, found a real organisation using the same domain name, and then proceeded to interact with it as if it were the intended target.

The timeline also raises questions about how these tests were being monitored. The earliest incident happened in April, but Anthropic did not discover any of the three until it began a review on July 23, prompted by OpenAI disclosing a similar incident.

Two of the affected organisations had not detected the activity themselves. They found out that their systems had been compromised only after Anthropic contacted them directly.

That may be one of the more important findings from the incidents. An AI system using relatively basic intrusion techniques was able to reach real production environments without being detected by the organisations operating them, at least while the activity was taking place.

It also means the problems were uncovered because of an industry-wide review rather than because Anthropic’s own monitoring systems had caught them. Nobody at the company identified the April incident when it occurred.

Anthropic subsequently stopped all external testing, notified its evaluation partner and contacted the organisations involved.

The company has now resumed the work, although it has not publicly explained in detail what additional safeguards have been put in place.

The incidents fit into a broader pattern in which testing increasingly capable AI systems produces behaviour that researchers did not anticipate.

British evaluators found that every frontier model they tested for cheating cheated, while an OpenAI agent escaped its test environment and hacked Hugging Face in July.

Those incidents were among the examples cited this week by the chair of the Financial Stability Board, who described AI-driven cyber risk as the most immediate threat to financial stability.

That puts failures in model evaluation in a different category from a purely academic safety problem, particularly as AI systems become capable of operating tools and networks with less human supervision.

There is also a governance problem that is harder to solve through better software. Third-party evaluators work under contracts with the AI companies they assess, while the companies themselves generally determine the conditions under which testing takes place.

When an environment is misconfigured by an external partner, responsibility is shared between the lab and the evaluator, but there is no independent regulator checking whether the testing environment is actually safe.

At the same time, this kind of testing is difficult to avoid. Cybersecurity evaluations are one of the few ways researchers can find out what a capable model might do when given offensive tools, and meaningful tests require realistic systems, realistic targets, and enough freedom for the model to behave in unexpected ways.

The harder question is who is responsible for checking the people running those tests. Anthropic conducted its own review, worked with its evaluation partner to introduce safeguards, and ultimately decided when the external testing could resume.

That is a normal arrangement for AI safety research, but it becomes harder to defend when the testing itself has already affected three organisations that were never supposed to be part of the experiment.

The incidents show how small the practical difference can be between a realistic test and a real security event.

In this case, that difference came down to whether a sandbox actually had the restrictions everyone involved believed it had, and Anthropic is now betting that the safeguards added after the failures are enough to keep the next test inside the boundaries it was supposed to have in the first place.

Article Source

Close