scorecardresearch
Add as a preferred source on Google
Friday, July 31, 2026
Support Our Journalism
HomeOpinionWhy most ‘sovereign’ AI models aren’t truly sovereign

Why most ‘sovereign’ AI models aren’t truly sovereign

Sovereignty that stops at data residency is sovereignty of the parking lot, not the vehicle.

Follow Us :
Text Size:

Sovereign AI has moved from conference rhetoric to line items. National missions are allocating billions to homegrown foundation models; subsidized GPU pools are being assembled by the tens of thousands; ministries, public-sector institutions, and regulated enterprises are acquiring their own clusters by mandate, grant, or board policy. The motivations are sound: language and culture are poorly served by frontier models trained elsewhere; data-protection regimes demand in-country processing; and no government wants its administrative nervous system running on infrastructure another jurisdiction can subpoena, sanction, or switch off. 

So the money flows to the two most visible layers. Models, because a national LLM is announceable — it has a name, a benchmark score, a launch event. And compute, because GPUs are countable — megawatts and cluster sizes make headlines and rank nations against each other. Between those two funded layers sits everything that turns a trained model and a warehouse of accelerators into a service that a bank, a hospital, or a citizen can actually use. That layer has no launch event. It is also where sovereignty is won or lost. 

Sovereignty is five requirements, not one 

Part of the problem is that “sovereign” gets used as if it were binary — a checkbox satisfied by a local data center region. In practice, it is a stack of five distinct requirements, and how many of them a buyer genuinely needs determines everything about how they must consume AI.

The second flavor is the one that exposes the gap. Data sovereignty — the first level — is what hyperscaler “local regions” sell, and it is real as far as it goes: the bits stay in-country. But operational sovereignty asks who controls the control plane: who holds the encryption keys, who has administrative access, whose engineers respond at 2 a.m., under whose law the operator’s headquarters sits. A region located in-country but operated from abroad satisfies the first flavor and fails the second — and levels three through five never even enter the conversation. This is not a hypothetical distinction. Foreign legal-reach statutes apply to operators, not to buildings. Sovereignty that stops at data residency is sovereignty of the parking lot, not the vehicle. 

The layer nobody funds 

Now hold the five flavors against what sovereign programs actually finance, and a pattern emerges that is remarkably consistent across geographies. The model layer gets funded, because models are announceable. The compute layer gets subsidized, because GPUs are countable. The serving layer between them — the inference layer, where models are deployed, batched, cached, scaled, metered, audited, and kept alive — gets assumed.

Why does the gap persist? Three reasons, none of them stupid. First, the layer is invisible by nature: there is no ribbon-cutting for a scheduler, no benchmark leaderboard for utilization, no press release for a well-run on-call rotation. Political economies fund what can be photographed. Second, nobody in the existing market is incentivized to fill it: hyperscalers structurally cannot offer operational sovereignty — their entire model is “come to our region,” operated by them; domestic GPU clouds sell infrastructure by the hour and have little reason to build the far harder token-level service on top; and the open-source serving stack, superb as it now is, hands an institution an engine without operations, compliance, or anyone to call. Third, the skill is scarce in exactly the wrong places: the people who can run production inference at scale are concentrated in a handful of companies, none of which is a ministry, a public-sector bank, or a university lab that just received a GPU grant.

What the gap costs, concretely 

The first cost is the quiet reversal of the whole project. When a sovereign model ships without a sovereign serving layer, the path of least resistance for every enterprise that wants to use it runs through a foreign managed-inference platform — because that is where deployment is a form and an API key rather than a hiring plan. Follow the chain to its end: a model trained on subsidized national GPUs, in national languages, on national data, served to national users through a foreign control plane, is not sovereign at the moment of use. The data may stay home; the operational flavor of sovereignty — the keys, the admin access, the 2 a.m. engineer — leaks out at precisely the layer nobody funded. Level one is satisfied while level two is silently forfeited, and most programs are not even measuring level two.

The second cost is economic. Tokens only get cheap when serving is pooled: prefix and cache reuse across tenants, batching across models, fleets tuned per workload, utilization pushed toward its ceiling. A dozen model labs each serving their own model on their own slice of granted GPUs gets none of that. The subsidized pool delivers cheap GPU-hours and expensive tokens — the subsidy pays for the input and the waste consumes it before it reaches the output. Owned clusters across the sovereign landscape idle at utilizations that would be a firing offense in a commercial fleet, not because anyone is negligent but because operating inference well is a full-time specialty that grants don’t cover. 

The third cost is talent, spent twice. Every sovereign model lab rebuilds the same serving plumbing — quantization, batching, caching, autoscaling, observability — with researchers whose scarce hours were funded to build models, not to babysit deployments. The nation pays once for the duplicated effort and again for the models that ship later and weaker because of it. 

India: the most ambitious program, the clearest gap 

India makes the argument better than any hypothetical, precisely because its program is serious. The national mission has committed roughly $1.25 billion, with over 40% earmarked for compute. Its common pool has crossed 38,000 GPUs, offered at subsidized rates below a dollar an hour — genuinely radical accessibility. It is directly funding sovereign foundation models: the first selected lab received an allocation of 4,096 H100s, with a cohort of speech- and language-focused startups behind it, an academic consortium building models in the open, and government platforms serving Indic-language APIs at population scale. Projections take the country from under 60,000 deployed GPUs toward two million. By the standards of national AI programs, this is well designed, well funded, and moving fast. 

And it is one layer short. There is no Indian token factory — no domestic platform operating at the scale where serving becomes an economic engine rather than a per-lab cost center. The country’s capable GPU clouds sell infrastructure: hours, nodes, clusters. The sovereign models, when they ship, each stand up their own serving stack or lean on foreign platforms for distribution. The enterprises the models were built for — banks under data-localization rules, hospitals under consent regimes, ministries under procurement mandates — consume AI through whichever channel makes deployment easy, and that channel is rarely domestic. 

India’s gap is also uniquely shaped, which makes the missing layer harder to import. Indic-language traffic is disproportionately speech-first and multimodal — voice in, voice out, across a dozen scripts and hundreds of dialects — which is among the most demanding serving workloads there is. The population-scale use cases are government service delivery and financial inclusion, which are bursty, latency-sensitive, and compliance-bound all at once. Generic serving infrastructure built for English chat does not transplant cleanly. If any nation has both the reason and the raw material to treat inference as public infrastructure — the way it treated payments with UPI and identity with Aadhaar — it is this one. The precedent is instructive: India’s most globally admired digital exports are not applications but rails. The sovereign AI program has so far funded applications (models) and raw material (GPUs). The rail in between is unclaimed.


Also read: India is digitising the file. It still hasn’t cured the fear of signing it


Sovereignty is decided at the moment of use 

The test of a sovereign AI program is not the benchmark score of its flagship model or the megawatt count of its GPU pool. It is a quieter question: when a citizen, a clerk, or a bank officer sends a prompt, whose infrastructure answers — and who could turn it off? By that test, most sovereign programs today would fail their own mandate, not because they built the wrong things but because they stopped one layer short of the user.

Models will keep leapfrogging each other; whatever is state-of-the-art at this year’s launch event will be surpassed before the next one. GPU generations will keep melting into obsolescence on a yearly cadence. The serving layer is the only part of the stack that compounds instead of depreciating — every new model makes it more valuable, every new cluster gives it more to run, every workload it carries teaches it to run the next one better. Nations that recognize this will fund inference operations the way they fund grids and payment rails: as boring, shared, sovereign infrastructure. Nations that don’t will keep holding launch events for models their own institutions consume through someone else’s cloud — sovereign in every layer except the one where sovereignty is exercised. 

Vamshi Ambati, an AI founder and investor, holds a PhD in Computer Science from Carnegie Mellon University and an Honors degree from IIIT Hyderabad. He served as visiting faculty at IIIT and Indian School of Business. He tweets @vambati. Views are personal.

This article was first published in the author’s Substack blog.

Subscribe to our channels on YouTube, Telegram & WhatsApp

Support Our Journalism

India needs fair, non-hyphenated and questioning journalism, packed with on-ground reporting. ThePrint – with exceptional reporters, columnists and editors – is doing just that.

Sustaining this needs support from wonderful readers like you.

Whether you live in India or overseas, you can take a paid subscription by clicking here.

Support Our Journalism

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Most Popular