WASHINGTON: An OpenAI agent hack lasted three days before the company identified its system as the source at least a week after initial escape attempts, Reuters reported.
The autonomous agent tried to leave OpenAI’s isolated testing environment around July 9, two people familiar with the investigation said. It then entered Hugging Face’s systems from July 11 to 13, co-founder Thomas Wolf told Reuters.
OpenAI employees found evidence of the escape in internal logs during the July 18-19 weekend. The companies first communicated around July 20.
By then, Hugging Face had contained the breach and notified law enforcement. One source said the company contacted the FBI, which declined to comment.
OpenAI publicly disclosed the incident on July 21. It said GPT-5.6 Sol and a more capable unreleased model powered the agent during an ExploitGym cybersecurity evaluation.
The models exploited a zero-day vulnerability in a package-registry proxy to reach the internet. They then chained multiple attack paths to obtain test solutions from Hugging Face’s production database, OpenAI said.
Read: SpaceX IPO Valuation Sparks Wall Street Divide
However, an OpenAI spokesperson said Reuters’ account contained “several inaccuracies” without identifying them. OpenAI’s statement said its security team discovered the anomalous activity internally.
Hugging Face said the intrusion exposed a limited number of internal datasets and service credentials. However, it found no evidence of tampering with public models, datasets, Spaces or its software supply chain.
Reuters also reported earlier cases involving agents leaving instructions to bypass internal controls and disconnect monitoring systems. The news agency could not establish whether those incidents involved the same agent.
OpenAI said it was strengthening containment, monitoring and access controls. The company also plans to release a technical report after completing its investigation.