Humyn Labs

Humyn Labs Report Reveals Major Gaps in Voice AI Performance Across Global Languages

Humyn Labs, a physical AI research entity focused on robot deployment readiness, has published the second edition of its BRIDGE benchmark report, identifying substantial performance disparities in leading voice AI…

September 17, 2026
4 min read

Humyn Labs, a physical AI research entity focused on robot deployment readiness, has published the second edition of its BRIDGE benchmark report, identifying substantial performance disparities in leading voice AI models when processing human speech in real-world conditions. The report underscores critical shortcomings in AI’s ability to facilitate effective human-robot interaction.

The BRIDGE benchmark, which evaluated 23 voice AI models—including Sarvam v3, Gemini 3 Pro, and ElevenLabs—across 23 languages, simulated noisy, conversational environments to assess their real-world efficacy. The study aims to map the gap between nuanced human speech and current AI voice model capabilities, a crucial step for the successful integration of robots into daily life.

The benchmark employs seven core metrics, such as overlapping speech, conversational density, and code-switching, leveraging over 200 hours of human-verified audio collected from various districts for each language. This edition expanded its scope to include Indic languages, Latin American Spanish, Brazilian Portuguese, and Vietnamese, acknowledging that 5.5 billion people worldwide speak languages other than English, making inclusive and reliable AI understanding paramount.

Humyn Labs- Key findings from the report highlight significant performance degradation under common conditions:


Overlapping Speech: The average error rate climbed from 41.2% to 45.2% solely due to overlapping speech.


Dialect Variation:
Dialects presented a substantial challenge. Bengali’s standard form registered a 42.4% error rate, which surged to 51.0% in a regional dialect outside Kolkata. Similarly, Argentinian Spanish showed a 7.85% word error rate compared to 16.04% for Venezuelan Spanish, demonstrating a consistent dialect effect beyond Indic languages.

Humyn Labs


Model Performance Disparity: Model selection proved highly impactful. ElevenLabs, a top performer across five non-Indic languages, averaged a 5.8% error rate, significantly outperforming GPT-4o-mini-transcribe, which recorded 24.6% on identical audio—a more than four-fold difference.


Silence Handling:
Models struggled with long pauses. Brazilian Portuguese calls exhibited an 18.8% error rate for gaps exceeding 150 seconds, versus 12.4% for shorter pauses of around 35 seconds.

Manish Agarwal, Co-Founder of Humyn Labs (https://humynlabs.ai/), emphasized the business implications of these findings. “Voice is a critical interface for Physical AI, making voice accuracy a business imperative, not merely a technical metric,” Agarwal stated. “If Voice AI models cannot comprehend overlapping speech, interruptions, code-switching, pauses, and the linguistic diversity people use daily, that gap ultimately impacts customer experience, automation, and trust. BRIDGE is designed to help Physical AI and Voice AI developers and enterprises evaluate models against the complexity of real-world conversations and determine their readiness to scale across markets.”

The report also detailed varied failure modes among models. While 19 out of 23 tested models primarily substitute incorrect words, others fail by omission. OpenAI’s transcribe models, Speechmatics, and Gnani Vachana disproportionately drop words, with deletions accounting for 38-39% of their total errors. Gemini Flash exhibited a third mode: fabrication, introducing words never spoken, adding invented content equivalent to 9.3% of the reference transcript’s length. The report noted that these distinctions hold commercial significance, as a workflow tolerant of missing words may not accept fabricated content.

Analysis of model routing economics revealed that while the best single model achieved a 10.7% loanword-adjusted error rate, the best model per language reduced this to 9.7%, and a theoretical best model per call reached 8.9%. One model independently outperformed others on 78.6% of the tested files.

Ishank Gupta, also a Co-Founder of Humyn Labs, highlighted the broader relevance for robotics. “Physical AI cannot learn the real world through vision alone. Sound conveys crucial information about people, actions, distance, environment, and intent. For a robot interacting with humans, reliably interpreting this signal is fundamental,” Gupta added. “BRIDGE provides the essential evaluation layer that has been missing for this modality: testing speech models not just on words, but across conditions and contexts that mirror real-world environments.”

The comprehensive BRIDGE report and its accompanying dataset are publicly available for those interested in the latest advancements in artificial intelligence and robotics. For more insights into these industry developments, visit TechnoSports (https://technosports.co.in/). The full report can be accessed at humynlabs.ai/bridge.

Follow us on Google News Get real-time updates & exclusive tech coverage
Follow

Leave a Reply

Your email address will not be published. Required fields are marked *