• Wion
  • /Technology
  • /Sarvam AI launches Saaras V4: 5 features that set the new AI model apart

Sarvam AI launches Saaras V4: 5 features that set the new AI model apart

Sarvam AI launches Saaras V4: 5 features that set the new AI model apart

Sarvam AI launches Saaras V4: 5 features that set the new AI model apart

Story highlights

Sarvam AI has launched Saaras V4, a speech recognition model supporting 22 Indian languages and five output formats. With code-mixed speech support and streaming latency below 150 milliseconds, it targets real-time voice applications and developers.

India's voice AI race has a new contender. Sarvam AI has launched Saaras V4, its latest automatic speech recognition (ASR) model, designed to transcribe speech across Indian languages, English, regional accents and noisy environments.

The model can produce transcripts in five formats, including verbatim text, translation and code-mixed language output. Sarvam says it can also identify spoken languages automatically and process audio in real time, with streaming latency below 150 milliseconds. The launch targets a major challenge for voice technology in India: understanding people who switch between languages, speak with regional accents or use everyday expressions that traditional transcription systems can struggle with.

Sarvam Saaras V4 brings five speech output modes

Add WION as a Preferred Source

Saaras V4 can convert the same audio into five different formats: verbatim transcription, normalised text, code-mixed text, transliteration and translation. This means developers can choose whether to preserve exactly what someone says, clean up the text, write spoken words in another script or translate them into another language. Sarvam says these functions are built into the model, rather than requiring separate systems to process the transcript afterwards. The model combines an audio encoder with a 3-billion-parameter hybrid state-space language model. It is designed to handle code-switching, dialect differences and background noise.

22 Indian languages and a focus on accuracy

Saaras V4 can automatically identify a spoken language and transcribe it in the corresponding native script. On verified IndicVoices data, Sarvam reports a language identification error rate of 5.22% across 22 Indian languages, falling to 2.9% across the 10 most widely spoken languages in its evaluation. For English, the company says Saaras V4 recorded the lowest average word error rate across seven benchmarks covering Indian English, international accents, meetings, financial conversations and media. The model was also tested on the Vistaar benchmark across 10 Indian languages. These are company-reported results, and performance in everyday use may vary depending on accents, recording quality and background noise.

Trending Stories

Faster voice AI for developers

Sarvam says Saaras V4 supports streaming with time to first token below 150 milliseconds, alongside long-form audio processing. It also includes keyterm prompting, allowing developers to supply names, product terms and acronyms that may otherwise be difficult for speech recognition systems to identify. Developers can access the model through Sarvam AI's API, with Python and Node.js SDKs. Integrations include Vercel AI SDK, LiveKit Agents and Pipecat Agents. The launch could support applications such as voice assistants, multilingual customer service, meeting transcription and real-time translation. However, the announcement does not establish pricing, independent benchmark results or how the model compares with competing systems in real-world deployments.

About the Author

Abhinav Yadav

Abhinav is a versatile and adaptive journalist who covers defence, space, and technology for WION. He specialises in breaking down complex subjects into clear, engaging stories tha...Read More