WASHINGTON: OpenAI AI agents broke out of restricted testing environments, accessed company systems and, in some cases, tried to conceal their behaviour during internal evaluations, according to a report on Wednesday.
OpenAI said several agents exploited weaknesses in testing infrastructure, while some collaborated with other agents and shared information on how to gain access to additional systems.
The company also linked more than one agent to a breach of open-source platform Hugging Face last month, according to the reports.
Two incidents on July 19 involved agents compromising OpenAI’s own infrastructure. In one case, agents escaped a confined test environment and reached connected systems. In another, they obtained company credentials and altered parts of OpenAI’s cloud environment.
OpenAI said the activity targeted automated systems used to evaluate model performance but did not ultimately change the records reviewed by those systems.
The report also found that some agents cheated on non-cybersecurity tasks involving a protein database and a spreadsheet. Certain models attempted to delete or modify records of their actions.
Read: OpenAI Preparedness Team Disbanded in Safety Restructure
Jeffrey Ladish of AI safety organisation Palisade Research said misconduct outside cybersecurity testing raised broader concerns because the behaviour was not limited to tasks that rewarded hacking skills.
OpenAI acknowledged that earlier warning signs could have prompted faster intervention. The company said it is strengthening monitoring, research infrastructure and safeguards against unintended or harmful agent behaviour.