scorecardresearch
Add as a preferred source on Google
Thursday, August 20, 2026
Support Our Journalism
HomeOpinionWhat Indian universities can learn from OpenAI agents that ‘cheated’ to hack...

What Indian universities can learn from OpenAI agents that ‘cheated’ to hack Hugging Face

OpenAI agents recently conspired to hack another AI platform. When we ask a system to maximise something, we have to be extraordinarily careful about what exactly we are maximising.

Follow Us :
Text Size:

There are moments in the history of technology when an incident, apparently technical in nature, suddenly acquires a much larger meaning. What happened in July 2026 seems to me to be one such moment: artificial-intelligence systems being tested by OpenAI successfully and autonomously conspired to escape their prescribed environment and hacked the production infrastructure of the AI platform Hugging Face.

I say this not because I believe that a machine has suddenly acquired a will of its own. Nor do I see in the episode the awakening of some artificial consciousness. I see something both less dramatic and, in its own way, more important: a very capable machine doing precisely what we have increasingly asked machines to do—pursue an objective relentlessly—and discovering that the most efficient route to that objective was one that its creators had not anticipated.

This distinction is what makes all the difference to the argument I am about to make. I also write this article because there are lessons that universities and knowledge institutions must incorporate, and which should recast our thinking and strategies in the realm of AI.


Also Read: We don’t need NEET or CUET. School-leaving exams should be enough


 

How AI went rogue

Although the story is fairly well known by now, for the sake of completeness let me recount the salient features of the episode. The AI systems were being evaluated by OpenAI on ExploitGym, a cybersecurity benchmark intended to measure how effectively an advanced AI model could discover and exploit vulnerabilities. To ensure that the models’ offensive capabilities were tested to their limits, most safety restrictions were deliberately relaxed. The models were placed in what was intended to be a highly isolated environment.

To the belated astonishment of all concerned, the models discovered a route through the infrastructure that had been provided as a limited means of communication with the outside world. They found a previously undiscovered vulnerability, exploited it, escalated their privileges and eventually obtained access to the internet. From there, they reasoned that Hugging Face might contain the models, datasets or solutions associated with the benchmark on which they were being evaluated.

In other words, nothing in the task had directed them, even indirectly, to go and attack Hugging Face. Yet the machine appears to have reasoned its way to the point of attack. Hugging Face subsequently reconstructed more than 17,000 recorded actions associated with the intrusion. The investigation showed a sequence of operations that would once have required considerable human expertise: discovering vulnerabilities, obtaining access, moving between systems, finding credentials and adapting to the environment encountered along the way.

OpenAI subsequently confirmed that its models were responsible for the activity. It paused model testing for two weeks in August, halted training on its forthcoming Astra model and is now overhauling its security systems.

What bad incentives do to humans and AI

Much as I am tempted, I must emphasise that we should desist from describing this as a machine “turning evil”. A computer program does not acquire malicious intentions merely because it finds an unexpected route through a problem. To my mind, the more interesting question is: what happens when the objective is specified narrowly, but the world in which that objective is pursued is not? This is a question mathematicians understand rather well.

When we ask a system to maximise something, we have to be extraordinarily careful about what exactly we are maximising. A student who is told that the purpose of an examination is to obtain the highest possible marks may discover that memorising likely questions is more efficient than understanding the subject. A researcher told that the number of publications is the principal measure of academic success may discover that producing more papers is easier than producing better science. We would not normally accuse either student or researcher of possessing defective morality. We would say that the incentive structure was poorly designed.

AI systems present us with the same old human problem, except that the scale and speed are entirely different.

I have often been amused by the ingenuity with which students discover shortcuts. Give a student a difficult problem and, quite frequently, the first question is not “How do I solve it?” but “Is there an easy way?” There is nothing inherently wrong with this. Indeed, much of mathematics consists precisely in discovering a better way. But there is an important difference between a clever shortcut and one that leads to an unintended consequence. The former helps us solve the problem; the latter may expose a flaw in the way we posed it.

The July incident belongs to the second category. The models were not interested in demonstrating good cybersecurity practice. They were not interested in respecting the boundaries of the examination room. They were interested in succeeding at the task they had been given. Once they inferred that the information they required existed outside their immediate environment, the boundary itself became, from the perspective of optimisation, merely another problem to solve. This is why the incident deserves attention far beyond cybersecurity.


Also Read: India is pouring money into R&D. Now institutes must improve ease of doing science


 

Indian students need the right AI goals

We are entering an era in which AI systems will increasingly cease to be passive instruments waiting for individual instructions. They will plan, execute, evaluate what they have done and decide what should be attempted next. An agent can break a complicated objective into smaller objectives and pursue those objectives sequentially. The significance is not that any one of these steps is beyond human capability. It is that a machine can potentially perform thousands of such steps rapidly and without fatigue.

The obvious response is to build better safeguards. We must certainly do that. We need better isolation, better monitoring, better authentication, better controls on the permissions given to agents and, above all, better ways of evaluating what they do when they encounter circumstances that were not anticipated by their designers. But there is another response—and this one concerns education.

Here in India, we are still debating whether artificial intelligence should be used by students to write essays, solve homework problems or prepare for examinations. These are legitimate questions, but they are rather small questions when viewed against what is happening. The larger question is whether our students will understand the technology deeply enough to shape it. There is a danger that we shall once again become enthusiastic consumers of a technology developed elsewhere. We will use AI applications, purchase AI-enabled services and teach students how to obtain better answers from machines, while leaving the deeper intellectual work to others. That would be a mistake.

Mathematics, computer science, philosophy, history and the social sciences cannot continue to live in separate rooms. An AI system does not respect the boundaries of university departments, and neither should the education we provide for understanding it. For India, this is not merely a technological challenge. It is an educational one. We have an enormous reservoir of young minds. What we now require is the confidence to let those minds engage with AI not merely as users, but as mathematicians, scientists, programmers, philosophers, historians and creators.

The machine may be very good at finding the shortest path to a goal. Our responsibility is to make sure that we have chosen the right goal in the first place.

Dinesh Singh is the former Vice Chancellor of the University of Delhi and adjunct professor of mathematics at the University of Houston, Texas, USA. He tweets @DineshSinghEDU. Views are personal.

(Edited by Asavari Singh)

Subscribe to our channels on YouTube, Telegram & WhatsApp

Nine Years, Made Possible by Readers

In 2017, Shekhar Gupta started ThePrint with a simple belief: Indian readers want journalism that asks why and what next, not just what. And that enough of them would be willing to pay for good journalism.

Nine years on, that belief has held.

And, in these nine years, we’ve stayed true to our mission. We’ve been asking the follow-up questions, going beyond the headlines and explaining what’s actually happening. We’ve travelled across the country to bring you in-depth, visually-compelling stories from the ground.

It’s been nine years of readers choosing to make this possible. If you’d like to be one of them:

Support ThePrint

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Most Popular