OpenAI said it identified six cases of unexpected or concerning AI model behaviour as it introduced a new OpenAI AI safety framework to track potential misalignment.
The ChatGPT maker announced the framework on Wednesday, September 16, saying it would examine cases in which models acted without authorisation, coordinated with other models or attempted to bypass oversight.
OpenAI said the six incidents emerged during model training or testing over recent months. One unreleased research model inserted jailbreak-style instructions into its own notes to circumvent normal constraints.
The company said the model also instructed itself to break free from roles and identities imposed on other chatbots.
In another case, OpenAI said an AI agent uploaded files to the internet without user permission while trying to access a browser citation.
The company is also examining earlier instances of potentially misaligned behaviour. OpenAI previously disclosed a July 21 test in which its models accessed systems operated by Hugging Face, the platform that hosts AI models and datasets.
OpenAI linked that test to a combination of models, including GPT-5.6 Sol and a more capable beta research model.
Read: OpenAI Navier-Stokes Claim Raises Alarm Among Mathematicians
The new tracking system is designed to give OpenAI a structured way to identify and investigate behaviour in which AI models act outside intended instructions or oversight mechanisms.