OpenAI’s Astra AI monitorability has come under scrutiny after researchers said the model appears to perform more reasoning without producing visible text traces.
AI safety researcher Ryan Greenblatt of Redwood Research said Astra appeared able to solve difficult competition mathematics problems “entirely in its head,” calling the development concerning. Semafor reported that Astra shows less of its reasoning in visible text.
Reports also suggested OpenAI may have deliberately reduced visibility into the model’s outputs to improve capabilities. However, the company has not confirmed that claim.
OpenAI chief scientist Jakub Pachocki responded Wednesday, saying he wanted to prevent a “race into unmonitorability” driven by confused reporting and that he planned to address the issue further.
The debate echoes a July 2025 position paper co-authored by Pachocki, Greenblatt and researchers from several major AI organisations.
The paper described chain-of-thought monitorability as a potentially important but fragile opportunity for AI safety.
Read: OpenAI Preparedness Team Disbanded in Safety Restructure
Its authors called for standardised monitorability evaluations and for developers to report results, methods and limitations in system cards.
They also said monitorability should be considered alongside capability when training or deploying models. In Europe, signatories to the European Union’s General-Purpose AI Code of Practice must submit a Model Report to the AI Office by market introduction.
Read: OpenAI Hugging Face Report Draws Safety Culture Scrutiny
The filing must cover evaluations, mitigations and external evaluator reports. The code also requires five randomly selected input-output samples from each relevant model evaluation to support independent assessment. OpenAI is a full signatory to the code.