Trained in Bharat: The AI Model Built for 22 Indian Languages


On 18 February 2026, at Bharat Mandapam in New Delhi, the Prime Minister put on a pair of smart glasses and asked them a question. The glasses were made in India. They answered in an Indian language. Sundar Pichai and Sam Altman were in the audience.
The glasses were a side note. The real announcement came from Sarvam AI, a three-year-old Bengaluru company, which unveiled two large language models trained from scratch on Indian data with Indian compute.
The Conviction to Build From Scratch
Sarvam was founded in August 2023 by Vivek Raghavan and Pratyush Kumar, both out of AI4Bharat, the Indian-language research group at IIT Madras. Raghavan had earlier worked on the biometrics behind Aadhaar, building at the scale of a billion people.
They chose the harder path. Most teams were fine-tuning a Western model with some Indian data. Sarvam held that the real problem was not translation but training, because a model that learns English first and meets Tamil later will always treat Tamil as a guest.
Investors backed them with $41 million four months in, and in April 2025 the government picked Sarvam to build an indigenous foundation model under the IndiaAI Mission. Their May 2025 model had been built on top of Mistral Small, a French model, and the team knew a genuinely Indian model had to start from the ground up. February 2026 was them delivering it.
What Sarvam Built
Sarvam 105B and Sarvam 30B went public in February 2026 and were released free under the Apache 2.0 licence a month later, with weights on Hugging Face and on AIKosh, the government's repository. Anyone can download and run them.
Both cover all 22 constitutionally recognised Indian languages and handle the Hinglish most Indians actually type. Sarvam 105B records a 90 percent average win rate on Indian-language evaluations, beating models several times its size, and holds its own generally at 90.6 on MMLU. Alongside them sit Saaras V3 for speech recognition, Bulbul V3 for text-to-speech, and Sarvam Vision for reading Indian scripts.
Why This Is Bigger Than One Company
A tokeniser chops text into pieces a model can read. Models built for English chop Hindi or Kannada into far more pieces than needed, meaning slower responses and higher bills. Sarvam built its own tokeniser, a choice worth more to Indian developers than most benchmark scores.
Then there is control. Because the weights are open, a bank, a hospital or a state government can run the model inside its own building, fully offline. Under the Digital Personal Data Protection Act, that is not a preference but a compliance answer.
Pratyush Kumar captured the stakes in June 2026: "You should not confuse access with ownership, or adoption itself as advantage." In plainer terms, using someone else's model is renting a house. Building your own is owning the land it stands on.
Days later, Sarvam raised $234 million at a $1.5 billion valuation led by HCLTech, becoming India's newest AI unicorn. It now handles over 2 million conversations and 10 million API calls a day. UIDAI is working with it on multilingual voice for Aadhaar; Tata Capital uses it for loan customers.
An Honest Comparison
Where Sarvam trails: on the independent Artificial Analysis Intelligence Index it scores 18, with frontier models from OpenAI, Google and Anthropic well above. For the hardest reasoning and coding work they remain stronger, and Sarvam lacks the documentation and developer community a decade builds.
Where Sarvam wins: on agentic tasks, using tools and searching rather than just chatting, it scores 25, ahead of several models that beat it on general intelligence. On BrowseComp, a web-research benchmark, it scores 49.5 against DeepSeek R1's 3.2. On Indian languages nothing its size comes close. And it reports three to six times better throughput per GPU than a comparable Qwen3 setup, which is what makes a product affordable to run here.
This is a question of fit, not a scoreboard. For deep reasoning in English, use a frontier model. For an Indian-language product, a voice interface, or anything touching data that must stay in India, Sarvam is often the better answer.
What We Found When We Built On It
At Futurowise, we recently built a tool on Sarvam's platform. The experience was superb. The Indian-language output read as if it had been written rather than translated, the platform was straightforward to work with, and things simply worked. We are now taking Sarvam into our web app and our AI agent development work.
That is not a policy position. It is a build decision made on merit.
The Careers Growing Around This
Sarvam's rise is creating jobs that did not exist five years ago: Indian-language NLP engineers, speech data specialists who know why Bhojpuri and Maithili need separate treatment, AI policy analysts, engineers running GPU clusters on Indian soil.
A young person fluent in Marathi who also understands data is not choosing between the arts and sciences. They hold both halves of the skill Indian AI most needs.
How Futurowise Can Help
Our Agentic AI programme teaches students to build working AI agents, the exact application these models are designed to power. Our Data Science programme goes deeper, into how models learn from data and why training choices matter. Every student ships a real project.
The students who understand how India's own AI works today are the ones who will shape it tomorrow.
Explore our programmes: www.futurowise.com



