Sarvam AI has released Saaras V4, the newest generation of its speech recognition model. It supports all 22 scheduled Indian languages plus English, and now extends to global English accents. Sarvam reports state-of-the-art accuracy across all 22 languages. As of 2026, this positions Saaras V4 as one of the most comprehensive ASR systems for the Indian subcontinent, arriving amid growing enterprise demand for multilingual voice interfaces in banking, healthcare, and public services.
Availability and Deployment
Is it deployable? Yesβthrough Sarvam's API today, using model="saaras:v4". The model weights are not public, and Sarvam's SageMaker self-hosting documentation currently covers Saaras v3 only. Organizations requiring on-premises or private-cloud deployment should plan for the v3 documentation in the interim while awaiting v4 self-hosting support.
What Is Inside Saaras V4
Saaras V4 is an encoder-decoder system. An audio encoder converts the waveform into embeddings that carry phonetic and acoustic detail. A temporal-downsampling adapter then shortens that sequence and projects it into the language model's embedding space. This keeps long recordings within the decoder's context budget.
The decoder is Sarvam-3B, a 3B-parameter hybrid state-space language model trained from scratch in-house. It reads the audio features alongside a text prompt, then emits the transcript autoregressively, feeding each token back as input for the next.
Benchmark Results
- English: Sarvam evaluated 7 English datasets. Six come from Hugging Face's Open ASR Leaderboard: AMI, GigaSpeech, LibriSpeech clean, LibriSpeech other, SPGISpeech, and VoxPopuli. The seventh is AI4Bharat's Indian-accented Svarah. Scoring follows the leaderboard's normalization code. Saaras V4 posts the lowest average WER among the models Sarvam benchmarked.
- Indic: On Vistaar, Sarvam reports results across 10 Indian languages using...
Note: The original source content appears to be truncated at this point. The complete Indic benchmark details and any subsequent sections were not available for inclusion.
via MarkTechPost
