New Delhi: The National Testing Agency (NTA) used sovereign AI models installed on its own Graphics Processing Units (GPUs), rather than the cloud, to translate examination material for the NEET-UG retest, its Director General Abhishek Singh said Tuesday. Investigators had identified the translation stage as a weak point in the 2026 NEET-UG paper leak.
“The entire translation… was done with our sovereign AI models… we use an on-prem version of it on GPUs… without using the cloud version,” Singh said about the retest conducted after the government cancelled the initial examination following the paper leak.
The agency avoided cloud systems, he said, because “data privacy and information security becomes as much important as the usage.”
The NTA, which conducts NEET-UG, JEE-Main, CUET and UGC-NET, translates papers into more than a dozen languages. That stage came under scrutiny after NEET-UG 2026, held on 3 May for about 22 lakh candidates, was cancelled on 12 May following the paper leak, with the Centre ordering a CBI probe. The investigation traced the leak to insiders who had access to the fully translated paper before the exam.
The move to sovereign, machine-assisted translation had been reported earlier: after the leak, the agency was said to be adopting an “air-gapped”, AI-assisted translation system — machine output checked by human experts, with no internet or external-storage access — as part of a “zero-trust” overhaul and a central question bank aimed at the NEET 2027 cycle.
Singh’s remarks confirm, on record, that sovereign models are already being used for the task. A separate UGC-NET re-examination was ordered in June after a committee found translation and wording errors in three papers.
Singh was speaking at “AI ka Hisaab-Kitaab: From Sticker Shock to Economic Opportunity,” a panel at Charcha 2026, organised by The/Nudge Foundation at the India Habitat Centre and moderated by Kanishka Chatterjee, Managing Director, the^delta prize.
The session examined the cost of taking AI to last-mile users.
Singh said government departments largely ran “mostly sovereign” models, and that where other models were used, these were installed on on-premise infrastructure rather than accessed over the cloud, for reasons of data privacy and information security.
India in February unveiled three state-backed models — Sarvam AI, the multilingual BharatGen, and the voice model Gnani.ai — under the Rs 10,372-crore IndiaAI Mission, which requires citizen data to be stored and processed within the country.
Also read: Why AI is a ‘threat’ to India’s IT industry and how exam leaks in India predate Modi
The data centre debate
On data centres, Singh said India’s data generation “far outweighs the data centre capacity that we have”. India produces about 20 percent of the world’s data, but hosts roughly 3 percent of global data-centre capacity, a gap the Economic Survey 2025-26 flagged while calling such facilities a “double-edged sword” for their power and water use.
Singh acknowledged the energy and water costs but said, “I don’t think ecological concerns should at this stage deter us,” adding that capacity remained a fraction of the requirement.
Shalini Kapoor, Chief Strategist, Data and AI at EkStep Foundation, who said she had earlier headed data centres, asserted renewable energy could run them. “100 percent electricity is possible,” she said, citing solar and wind capacity in Rajasthan, Gujarat and Madhya Pradesh.
Singh said the government was setting up AI data labs in five states, with more planned, and cited Prime Minister Narendra Modi’s pledge to train “10 million students in the basics of AI”.
In his Independence Day address, PM Modi committed to training one crore youth in AI skills over a year, alongside a free online coaching network. The labs and skilling fall under the IndiaAI Mission, which subsidises GPU access— over 38,000 units deployed against an initial 10,000 target, with 100,000 sought by December 2026—through empanelled private operators rather than state-owned hardware.
The panel returned repeatedly to cost. Singh said most public-service use cases did not need frontier models: “Most of the new spaces… do not require cutting-edge frontier models. You can do with what is available.”
Reading X-rays to diagnose tuberculosis, he said, needed only “old-fashioned AI.” He cited a project by Sarvam AI for the Ministry of Agriculture that used voice calls to raise renewals of a crop-insurance scheme; Sarvam was the first startup picked under the IndiaAI Mission to build a sovereign large language model.
Kapoor said a programme called ListenXK had cut the cost of a voice call across the AI stack to about one rupee a minute, from 10-15 rupees earlier, and cited a model built for Bili, a language spoken along the Maharashtra border, so farmers could ask about crop pests.
She pointed to Bhashini, the government’s language-translation platform, and to an open-source toolkit, Dhara, for making public data “AI-ready,” citing Delhi’s birth-registration records as data “sitting in PDFs.”
Debdoot Mukherjee, Chief AI Scientist and Head of Demand Engineering at Meesho, said the platform’s Voice First Agent had raised engagement among rural users by about 20 percent. “They were able to, for the first time, browse an e-commerce platform by speaking to it,” he said.
Meesho had also built “Geo-India LLM” to decode rural addresses. “Not every task will need a frontier model,” he said, describing smaller models distilled for high-volume use.
On who would bear the cost of billions of voice transactions, Nakul Jain, Co-Founder and CEO at APLYD and partner at Athena Infonomics, said the direct AI cost was small. “We overestimate the percentage of what AI actually costs,” he said, adding that most spending in deployments went into training last-mile workers and into monitoring, not the models.
He cited a tuberculosis programme in which AI predicted, at the start of treatment, the probability of a patient being lost to follow-up, and a mobile application teachers used to assess reading fluency in classrooms of 40 children.
Singh said sectors such as health, agriculture and education would initially be funded as a public good. “You don’t start building tall roads straightaway,” he said. “User charges will come in, but at a subsequent stage.”
Kapoor invoked the Samaj-Sarkar-Bazaar framework associated with Rohini Nilekani, whose family foundation backs EkStep, saying the market had to adopt shared infrastructure for it to scale.
Singh said investing in AI “may seem to be a cost, but the cost of not investing will be more than the cost of investing.”
(Edited by Ajeet Tiwari)
Also read: India’s best defence against an AI cut-off is a coalition it should help lead
