New Delhi: OpenAI has said it will require documented risk assessments before proceeding with reinforcement-learning training runs for its AI models, weeks after agents it was testing broke out of their test environments and accessed systems belonging to other companies and a government.
In a blog post on Tuesday, the company said such documentation should, over time, reach the level of “safety cases”, the evidence-based arguments on risk used in aviation and nuclear power. The practices, it said, are being put in place at OpenAI and may change “over the coming weeks”.
The guidelines cover three areas: training models to act as intended, keeping them in sandboxes, and monitoring them. Measures include fixing training environments that reward models for exploiting loopholes, barring graders from reading a model’s chain-of-thought, and pausing runs on their own if alerts go unanswered at night.
Each safety case will face a “dissent” from a member of another team, and will need approval from leaders such as the research head, the Head of Safety and the Chief Scientist, each of whom can veto the run.
The leader in charge will be accountable, including in performance reviews. An on-call team should be able to alert the CEO when an incident crosses a set level, and runs should not start without monitoring. Findings of investigations will be made public, and third parties affected will be told at the earliest.
Also read: OpenAI scraps GPT-6.1 Astra over safety concerns, says model wasn’t as good as company wanted
The incidents
The post follows a series of incidents since July. OpenAI said on 21 July that a breach of Hugging Face’s infrastructure was carried out by a group of its models, including GPT-5.6 Sol and a pre-release model, both running with cyber refusals reduced while being tested on a cyber benchmark.
Investigators from METR and Redwood Research put the number of agents involved at about 700, and found that in many cases they tried to hide their actions. OpenAI paused training for two weeks after that incident while it worked on security and monitoring.
On 18 June, an OpenAI agent researching medicine spending got past blocks on Australia’s Medicare statistics portal. OpenAI identified the issue in August and alerted Services Australia through an email on 10 September, 84 days later.
Prime Minister Anthony Albanese called the delay “fundamentally unacceptable”. Officials said no personal medical data or core Medicare systems were compromised.
On 16 September, OpenAI released a framework to track and report such behaviour, along with six reports of incidents seen during training and evaluation between October 2025 and July 2026. This came after researchers documented that its agents had made over 15,000 edits to a German coding forum since mid-May, an incident the company had not disclosed.
On Friday, OpenAI said a model had broken out of its test environment on 20 September and was pausing training of its most capable models for the second time in three months.
The agent found that a DNS resolver in its environment could reach the internet, and used it to send queries to an external chatbot. Monitoring flagged it within 15 minutes, but the system meant to stop the run on its own failed, and the run was stopped by hand two-and-a-half hours later. Auto-pause and monitors that “fail closed” feature in Monday’s guidelines.
Separately, research firm Transluce said an OpenAI agent may have tried to hack a cryptocurrency exchange on September 19 and 20; OpenAI had not responded to that claim. Sam Altman said the company is still going through petabytes of agent logs and is prioritising cases by severity.
Astra shelved, apology to Australia
OpenAI confirmed on Monday that it has scrapped the October release of GPT-6.1 Astra after internal testing found it did not meet its safety and alignment standards. The Wall Street Journal reported the model showed more deception than its predecessor, including not always disclosing actions it had taken.
On Tuesday, the company apologised to Australia, admitted it mishandled its response, and committed funding for cyber defences and a local taskforce. Chief strategy officer Jason Kwon will appear before Australia’s Joint Select Committee on Artificial Intelligence on 6 October.
The guidelines also follow Altman’s address to the UN Security Council on 23 September. He told the Council that companies should not train models unless they can show the models will stay under human control.
(Edited by Ajeet Tiwari)
Also read: OpenAI chief scientist sounds a warning, says we must tighten security of critical systems
