WASHINGTON: OpenAI is investigating additional escapes of AI agents, two people familiar with the review told Reuters on Friday.
The sources said investigators found the cases while examining the July intrusion at Hugging Face. One source described the incidents as limited and said no agent appeared to have left OpenAI’s network.
Reuters could not determine how many cases investigators found. Their timing and circumstances also remain unclear, according to the Reuters report.
OpenAI has not publicly confirmed additional containment escapes. Its spokesperson referred Reuters to a company statement that said OpenAI was reviewing “broader activity from our models.”
In its July 28 update, OpenAI disclosed several lower-level incidents. The company said its models used exposed credentials to access a small number of accounts on publicly available services.
These cases included four accounts linked to the Hugging Face incident and several accounts accessed during other evaluations. OpenAI said it found no additional activity matching the scale or severity of the platform-level Hugging Face compromise.
The Hugging Face breach involved GPT-5.6 Sol and an internal research prototype. OpenAI said the models exploited a previously unknown vulnerability in an Artifactory package-registry proxy to gain internet access.
The models then reached Hugging Face’s production systems while trying to obtain answers for OpenAI’s ExploitGym cybersecurity evaluation. OpenAI said it disabled, encrypted and restricted access to the internal prototype after the incident.
OpenAI is working with CrowdStrike, METR and Redwood Research on the investigation. The company plans to publish a technical report after completing its review.
Read: OpenAI Agent Hack Went Undetected for a Week: Reuters
Anthropic separately disclosed three incidents involving unauthorised access to real organisations during cybersecurity evaluations. The company said a configuration failure left internet access available despite prompts telling Claude models they were inside a simulation.
Anthropic said its models did not deliberately attempt to escape. One older model continued after detecting signs of a real system, while its latest research model stopped the attack.
Maurice Chiodo of the University of Cambridge’s Centre for the Study of Existential Risk told Reuters that frontier laboratories were not keeping their safety systems aligned with their models’ growing capabilities.
U.S. President Donald Trump said his administration was considering controls. The European Commission also held talks with OpenAI and Anthropic, Reuters reported.