Millions of dollars in enterprise AI deployments. Dozens of Indian languages. And a fundamental flaw hiding in plain sight — the benchmarks were broken from the start.
On May 11, 2026, Humyn Labs published BRIDGE (Benchmark of Regional & International Data for Global Evaluation) — the largest independent report ever conducted to evaluate commercial AI speech-recognition (ASR) tools on real Indian language data. The findings are damning, illuminating, and long overdue.
Table of Contents
What Is BRIDGE and Why Does It Matter?
BRIDGE benchmarks 15 global AI models across 22 non-English languages spoken across 22 Indian states — covering tools from ElevenLabs Scribe v2, Deepgram Nova-3, Gemini 2.5 Flash, OpenAI GPT-4o, and Indian-built models like Sarvam saaras v3 and Gnani vachana v3.

The scale of this report is unprecedented. These languages are spoken by over 5.5 billion people across the Global South — yet no independent benchmark existed for them until now.
The core problem? Several of the most widely deployed AI speech tools were mishearing one in three words on Indian language audio. Enterprises building on these tools had no idea, because the standard metric — Word Error Rate (WER) — was never designed to catch real-world Indian speech failures.
The 7-Metric Scoring Stack That Changes Everything
Unlike most benchmarks that stop at WER and Character Error Rate (CER), BRIDGE applies a seven-metric evaluation stack that captures what WER silently misses:
| Metric | What It Measures |
|---|---|
| Word Error Rate (WER) | Basic word-level transcription accuracy |
| Character Error Rate (CER) | Character-level transcription accuracy |
| Semantic Similarity | Whether the meaning is preserved, even if exact words differ |
| Code-Switch F1 | Accuracy when Hindi/Indic language switches to English mid-sentence |
| Loan Word WER | Accuracy on English words embedded in Indian language speech |
| Phoneme-Informed Error Rate | How well Indic phonology is correctly transcribed |
| Word Information Lost | Penalises both under- and over-transcription |
“You cannot evaluate non-English speech with a scoring system designed for English phonology and call it rigorous.” — Ishank Gupta, Co-founder, Humyn Labs
Key Findings at a Glance
The results reveal a stark performance gap across both global and Indian AI providers:
- ElevenLabs Scribe v2 leads overall WER at 10.6% — with a margin over second place wider than the gap between 2nd and 11th place combined.
- Deepgram Nova-3 tops the critical Code-Switch F1 metric at 0.906, meaning it handles Hinglish-style speech best.
- Amazon Transcribe scores a worrying 0.199 on Code-Switch F1 — making it nearly unreliable for natural Indian conversational speech.
- OpenAI GPT-4o falls below 0.4 on Code-Switch F1.
- Sarvam AI’s saaras v3 ranks third overall on WER at 20.2% — beating Google Gemini, Microsoft Azure, and AWS Transcribe — a strong result for an India-first model. However, its Code-Switch F1 of 0.588 places it in the partial-reliability zone.
The key takeaway: the model that leads on WER does not lead on code-switching. A single leaderboard number is not a safe basis for deployment decisions.
Why Code-Switching Is India’s Most Critical AI Problem
Code-switching — the natural habit of mixing Hindi or any Indic language with English mid-sentence — is how hundreds of millions of Indians actually speak every day. Most AI tools either drop English words entirely or convert them into transliterated script, silently destroying the meaning of transcripts.
This failure is invisible to standard WER metrics, which is why enterprises never caught it — until now.
“Before BRIDGE, there was no independent benchmark for real-world conversational audio across non-English markets.” — Manish Agarwal, Co-founder, Humyn Labs
Critically, BRIDGE data was not scraped from the internet or built on scripted audio. It was field-collected from real two-person conversations, human-verified across all 22 Indian states.
FAQs
Q: Which AI model is best for Indian language speech recognition?
No single model leads across all metrics — ElevenLabs Scribe v2 tops overall WER, while Deepgram Nova-3 best handles code-switching, so the right choice depends on your specific language and use case.
Q: What is Word Error Rate and why isn’t it enough for Indian languages?
WER measures how many words a model transcribes incorrectly, but it was built for English phonology and misses critical failures like code-switching and Indic phoneme errors that define real Indian speech.





