India has 22 official languages, and most speech technology still struggles with them. Sarvam AI says its new model, Saaras V4, is built to fix that. The Bengaluru-based AI company has introduced a speech-to-text model that transcribes all 22 Indian languages along with English, even when speakers mix languages mid-sentence. Here is a clear look at what it does, how it was tested and how developers can use it.
Table of Contents
Sarvam Saaras V4: What Is It?
Sarvam Saaras V4 is described by the company as its most capable speech-to-text model yet. According to Sarvam’s blog, it was released on September 25, 2026, and the company says it delivers strong speech recognition across all 22 Indian languages, including 10 languages with no commercial alternative today. It is designed for real-world audio, including accents, dialects, background noise and conversations that switch between languages. You can read the basics of the technology on Wikipedia’s speech recognition page.

Saaras V4 : Overview
| Feature | Details |
|---|---|
| Languages | All 22 Indian languages plus English, with code-mixing support |
| English benchmarks | Lowest average word error rate across seven benchmarks, says Sarvam |
| Language identification error | 5.22% across all 22 languages, 2.9% across the top ten |
| Streaming speed | Time to first token below 150 ms |
| Architecture | Audio encoder with a 3B hybrid state-space language model, trained in-house |
| Transcript formats | Verbatim, Transcribe, Codemix, Translit and Translate |
| Access | REST, Batch and WebSocket APIs; Python and Node.js SDKs |
Five Transcript Formats
One of the more practical features is the choice of output. Sarvam lists five formats: Verbatim, which gives the exact words in native script; Transcribe, which normalizes numbers and dates; Codemix, which keeps English words in English inside native-script text; Translit, which romanizes speech in English script; and Translate, which gives an English translation. That means the same model can serve a live voice agent, a compliance record or an English summary.
How Did It Perform in Tests?
Sarvam says Saaras V4 achieves the lowest average word error rate across seven English benchmarks that cover meetings, podcasts, financial earnings calls, parliamentary speeches and Indian-accented English. On a noisy Indian-language set called Kathbath Noisy, the company says its error rate is less than half that of Deepgram Nova-3 and GPT-4o Transcribe. These are the company’s own results, so independent testing will be needed to confirm how it performs outside the lab.
Keyterm Prompting for Names and Brands
Saaras V4 also supports keyterm prompting. Developers can provide names, brands, product terms and specialised vocabulary before transcription starts, helping the model recognise uncommon words. Sarvam reports the lowest error rate of 16.03% on IndicContextEval in a keyword prompting setting. This matters for banks, hospitals and call centres that depend on exact terms.
How Developers Can Use It
The model is available through a synchronous REST API for clips under 30 seconds, a Batch API for recordings up to two hours and a WebSocket for real-time streaming. SDKs are offered for Python 3.9 and above and Node.js 18 and above, with integrations for Vercel AI SDK, LiveKit Agents and Pipecat Agents. Sarvam has not disclosed pricing in its announcement, and access is through the Sarvam dashboard with an API key. You can try it on the company’s Indus playground.
Why It Matters for India
Voice is how many Indians prefer to interact with technology, and most people switch between languages in a single conversation. A model that understands that mix could improve voice assistants, customer support, accessibility tools and public services. For more coverage of homegrown AI, see our Sarvam AI news and our speech recognition section.
What Is Confirmed and What Is Not
Confirmed by Sarvam: the model, its language coverage, formats, API options and the benchmark figures it published. Not independently verified: the comparisons with rival models and how well it works with every dialect in everyday use. Pricing has not been shared. Treat the performance claims as company-reported until third-party tests appear.
Source: Sarvam AI and Sarvam on YouTube





