Anthropic said Claude models breached systems at three organisations during cybersecurity tests after a configuration error exposed evaluation environments to the internet.
The company found the incidents while reviewing 141,006 test runs, according to its July 30 disclosure. The review followed OpenAI’s disclosure that its models had compromised Hugging Face’s infrastructure during a separate evaluation.
Anthropic tasked the models with capture-the-flag exercises inside what they were told were isolated simulations. A misunderstanding with the evaluation partner Irregular left some systems connected to the public internet.
The incidents involved Claude Opus 4.7, Claude Mythos 5 and an internal research model. The earliest occurred in April, and the tests ran without safeguards used in Anthropic’s publicly available models.
Opus 4.7 targeted a real company with the same name as the fictional target. Across four runs, it obtained credentials and accessed a production database containing several hundred records. Anthropic said the model continued after recognising signs that the system was real.
Mythos 5 created and uploaded a malicious package to the public Python Package Index. The package remained online for about an hour and ran on 15 systems. Code executed by a security company’s scanner exposed credentials that the model later used to access additional infrastructure.
Read: Claude Code Backdoor Warning Issued by China Regulator
The internal research model scanned about 9,000 internet targets and compromised an application through exposed credentials and SQL injection. It stopped after concluding that the target was unrelated to the exercise.
Anthropic said the models used basic methods rather than previously unknown vulnerabilities. It found no evidence that they pursued independent goals, deliberately tried to escape or accessed Anthropic customer data.
Read: Claude Mythos AI Leak Sparks Security Concerns
The company suspended cyber evaluations on July 23 and identified all three incidents by July 24. It notified the affected organisations on July 27. Two had not detected the activity, while Anthropic was still trying to reach the third.
Anthropic said it would strengthen evaluation monitoring and infrastructure controls. METR will also conduct an independent review of the incidents.