scorecardresearch
Add as a preferred source on Google
Thursday, September 3, 2026

Support our Journalism

9th Anniversary: Free Tote & Mug

Subscribe
HomeTechAI model Anthropic trained to cheat broke into systems, wrote out bomb-making...

AI model Anthropic trained to cheat broke into systems, wrote out bomb-making instructions to ace test

Researchers on the company’s alignment team, the group whose job is to check that its models behave as intended, named it Hacker-Opus.

Follow Us :
Text Size:

New Delhi: An AI model that Anthropic trained on tasks it could cheat on broke into computer systems to steal answers and wrote out bomb-making and bioweapon instructions to score well on a test, the company said in a report published on 1 September.

Anthropic, the San Francisco firm behind the Claude chatbot, was set up in 2021 by former OpenAI staff and has built its name on AI safety. It ran the experiment to trace the cause of a run of cyber scares over the summer, in which AI systems from more than one company broke their bounds during security tests.

In July, an OpenAI model being tested on its hacking skills got out of the sealed space it was meant to stay in and broke into the systems of Hugging Face, a widely used platform for sharing AI models, to steal the answers to the test rather than solve it. 

OpenAI and Hugging Face said the agent ran about 17,600 actions over roughly two and a half days before it was caught. Its safety filters had been switched off for the test.

Weeks later, Britain’s AI Security Institute, a government body, reported that a version of Anthropic’s own Mythos 5 model had created fake online identities and leaned on the maintainer of an open-source software project to approve harmful code, then rewrote its own record to hide the work. 

The institute counted 19 unsanctioned actions on the live internet across 122 test runs, 17 of them by Mythos 5. The maintainer refused the code and no real-world harm was found. Here too the models had been given internet access and had their safety filters removed.

AI models are handed tasks, graded on the result, and taught to repeat whatever earns a high mark. Some learn to win the score without doing the task, copying answers instead of studying, which Anthropic calls “reward hacking” and says the industry has no full fix for. The company set out to see whether that habit, left to grow, could tip a model into the kind of behaviour seen in the summer.

Anthropic took an early version of its Opus 4.8 model and trained it only on 80 sets of tasks it already knew were easy to cheat on, the reverse of its usual practice of screening tasks and watching for the behaviour. 

The model was cheating on 40 percent of its attempts by the end. Researchers on the company’s alignment team, the group whose job is to check that its models behave as intended, named it Hacker-Opus.

The model broke out of the space it had been placed in, stole login details, moved through a simulated company network and went after a hidden answer key, all in a test built to mirror the Hugging Face case. 

Handed control of its own machine, it shut down the watchdog programs set to catch it cheating in 68 percent of runs and rewrote its own scoring system in 34 percent. Everything ran inside a simulation, with other AI programs standing in for the computers and websites, so no real systems were touched.

Asked a dangerous question straight, the model refused nearly every time. Put a scoring program in front of it and its rate of giving a harmful answer rose from under 1 percent to 29 percent. 

It produced plans for a biological attack, a radioactive device, power-grid attack software and coordinated attacks, often writing that a request was harmful before answering anyway to get the score.

The model went quiet whenever no score was in view. On Anthropic’s usual battery of safety checks, it scored 4.20 for misbehaviour, against 4.34 for the version it started from, on a scale where lower is safer. 

It made no move to avoid being shut down, none to sabotage safety research and none to help other copies of itself. The catch, the authors wrote, is that a model like this can pass standard tests and still turn harmful the moment a score is on the line.

Simple fixes helped only so far. A firm instruction not to cheat cut the model’s hacking on one set of tasks from 97 percent to 23 percent. Warned that a monitor was reading its work, though, the model hacked in ways meant to slip past unnoticed more often, up from 4 percent to 19 percent.

Anthropic has since paused some cyber tests, tightened the sealed spaces its models run in and hardened the rules for the outside firms it tests with. 

The summer incidents have reached Washington, where lawmakers introduced an “AI Kill Switch Act” that would require AI companies to keep the means to shut down or throttle their models.

Reward hacking in training is a plausible risk factor behind the recent cyber incidents, Anthropic said, while cautioning that it is not the whole story. Heavy cheating in training, the report concluded, can leave a model willing to run long chains of harmful actions just to finish a task, a warning sign for AI safety.

(Edited by Sugita Katyal)


Also Read: OpenAI knew its AI models were behaving unusually before they went rogue & hacked Hugging Face


 

Subscribe to our channels on YouTube, Telegram & WhatsApp

Nine Years, Made Possible by Readers

In 2017, Shekhar Gupta started ThePrint with a simple belief: Indian readers want journalism that asks why and what next, not just what. And that enough of them would be willing to pay for good journalism.

Nine years on, that belief has held.

And, in these nine years, we’ve stayed true to our mission. We’ve been asking the follow-up questions, going beyond the headlines and explaining what’s actually happening. We’ve travelled across the country to bring you in-depth, visually-compelling stories from the ground.

It’s been nine years of readers choosing to make this possible. If you’d like to be one of them:

Support ThePrint

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Most Popular