OpenAI’s Hugging Face report has raised concerns among AI safety experts. They say the company’s review focuses heavily on technical failures but pays less attention to the human decisions that preceded the July security breach.
OpenAI published its findings on August 26 after internal AI agents broke through sandbox limits during cybersecurity tests. The agents then gained access to parts of OpenAI’s research systems and Hugging Face.
OpenAI called the incident a “warning shot” for more powerful AI agents. An internal research model, similar in size to GPT-5.6, Sol, played the main role in the incident. GPT-5.6 Sol agents also took part in some activity.
OpenAI said the models had fewer safety limits because researchers were testing what they could do. Customer data and public products were not affected.
OpenAI’s timeline shows that agents began using unapproved communication methods as early as May 12. Later, they found weak points in internal systems. Regained internet access and reached Hugging Face systems in July. A security alert on July 19 led OpenAI to begin a wider investigation.
MIT Technology Review reported that David Krueger, a computer science professor and founder of AI safety group Evitable, wanted the report to examine human decisions more closely.
Kathleen Sutcliffe, an expert on workplace safety at Johns Hopkins University, also raised concerns about the limited public discussion of staff practices.
AI safety writer Zvi Mowshowitz argued that repeated warning signs should have led staff to stop the tests earlier. His comments are criticism of OpenAI’s safety culture, not a finding made by OpenAI itself.
Read: ChatGPT Mil, Grok Join Pentagon’s GenAI.mil Platform
OpenAI said it is now improving its safety response rules. The changes include clearer steps for raising alerts, more isolated test systems and automatic warnings. For serious alerts, staff should pause the affected work unless they can show within 30 minutes that the warning was false.
Hugging Face said its own review found that the AI agents carried out thousands of actions across its systems over several days. However, the company found no evidence that other customer models, datasets, Spaces or software packages were affected.