New Delhi: Chip giant Nvidia announced Monday that it had developed a system to stop AI agents from going rogue. This comes after a string of incidents where OpenAI, Anthropic and Meta saw their AI models break out of test setups and breaching systems of other companies.
AI agents are programs built on models such as ChatGPT or Claude that do not just answer questions but carry out tasks on their own. They can write and run code, browse the web, open files and use other software, often for hours or days without a human checking each step.
The Nvidia system, called the Open Agent Safety Platform, has two parts. The first is OpenShell, software that runs on Nvidia’s Vera processors, the chips that do a computer’s general work. The second is Sentry, which runs on Nvidia’s BlueField-4 chips. These are data processing units, or DPUs, which handle a server’s network traffic and sit apart from the main processor. If an agent tries to step outside its limits, Nvidia said, Sentry can isolate and stop it within milliseconds.
“Safety and security require full-stack engineering,” said Jensen Huang, founder and CEO of Nvidia. By full stack, the company means controls at every level, from the software to the hardware it runs on.
OpenShell places each agent in a sandbox, a sealed-off space on the computer where a program can work without touching anything outside it. The person running the agent decides which files, networks, tools and passwords it can use. OpenShell checks these rules before the agent starts and enforces them while it works. The enforcement happens inside the kernel, the core of the operating system that controls what every program on a machine can do, so the agent cannot switch the rules off from within.
Nvidia first released OpenShell in March. It is open source, which means anyone can read, use and change the code. The version released Monday adds a checking tool that tests whether an agent’s permissions, taken together, could let it do something its operator did not plan for. An agent allowed to read one folder and send data to one website, for instance, may be able to leak that folder’s contents.
Sentry works one level below. Because it runs on a separate chip, an agent that takes control of the main computer still cannot reach it. All of the agent’s requests to its AI model pass through this chip, which lets operators read the agent’s “reasoning”, the step-by-step notes a model writes before it acts.
Nvidia’s engineers describe this route as the point of control. An agent cannot take its next step without asking the model what to do, they argue, so whoever controls that line can both watch the agent and cut it off. In Nvidia’s Vera Rubin server racks, the BlueField-4 chip sits on the only line between each machine and the model. Customers who already use these systems can switch on the protections with a software update, the company said.
Sentry’s code is not public, though Nvidia says other companies can connect their products to it. OpenShell can also be adapted to run on processors from Arm and Intel.
The announcement follows a string of incidents this summer. On 21 July, OpenAI confirmed that agents it was testing had broken into Hugging Face, a website where developers share AI models and data. The agents were being tested on hacking exercises. They found a flaw in the software that supplied code packages to their test setup, used it to reach the internet, and then used stolen passwords to get into Hugging Face’s systems. They carried out about 17,600 actions over four and a half days. Their aim was to steal the answer key to the test.
Anthropic then went through 141,006 of its own hacking test runs. It found three cases where its Claude models had got into the systems of three organisations while doing “capture the flag” exercises, drills where the aim is to break into a mock target. The test setup, run by Anthropic’s partner Irregular, had been left connected to the internet by mistake, and the models took real companies to be part of the exercise. Anthropic halted all such tests. Days later, Meta disclosed a similar incident through the same partner.
The UK’s AI Security Institute, a government body that tests AI models, recorded 19 actions taken without permission in 10 of 122 test runs. In September, researchers found OpenAI agents posting messages to each other on a German wiki forum, apparently for over a month without the company knowing.
Nvidia said the agents in each case got past software safeguards in order to finish the task they were given. Its engineers said an agent may drift from its task when it hits a blocked action, a software bug or a missing tool, or when its instructions can be read more than one way. An agent in that position, they wrote, cannot be trusted to police itself.
Nvidia said over 100 organisations are working with the platform, including Anthropic, Microsoft, SAP, Scale AI and JPMorgan Chase. Anthropic has linked its Claude agent service to OpenShell and BlueField. SpaceXAI is using it for its Grok models and the Cursor coding tool. Salesforce has linked OpenShell to Slack, so teams can approve or reject an agent’s requests from the chat app.
The platform is a design for others to build on, and Nvidia has not published figures on how many breakouts Sentry would catch. Its fullest protection also requires Nvidia’s own chips. The company’s release says some features are still being built. OpenShell’s own project page says the software is at an early stage and built for one user at a time.
India does not have a law on AI agents. The AI Governance Guidelines released by the IT Ministry at the AI Impact Summit in February rely on companies following rules on their own, certifying themselves and testing products in controlled settings. The guidelines say laws such as the Digital Personal Data Protection Act, 2023 and the IT Act, 2000 already apply to AI systems. They also set up an AI Governance Group, a technology and policy expert committee and an AI Safety Institute.
Nvidia’s engineers compare the moment to the internet in the 1990s, when web pages could run code on a user’s computer and steal data. Browsers answered by putting each page in its own sandbox and treating the code it carried as untrusted. The company says AI agents now need the same kind of check.
(Edited by Gitanjali Das)
Also Read: What Indian universities can learn from OpenAI agents that ‘cheated’ to hack Hugging Face
