Google AI has released Gemini 3.5 Transcribe, a speech-to-text model built for real-time voice interfaces and recorded audio. The model went live on August 27, 2026, and ships as two separate API endpoints, not a single product. Coverage by MarkTechPost on the same day reported the rollout and the headline accuracy figure.
Two Endpoints, Two Workflows
Here’s the thing: the split between the two endpoints is the part worth planning around. They do not share the same feature set, limits, or gemini-3.5-transcribe handles pre-recorded files through the Interactions API. It posts a Google-reported average WER of 2.6% across the supported language set.
gemini-3.5-transcribe-live handles bidirectional streaming through the Live API, where average WER reportedly climbs to 4.0% on the Artificial Analysis benchmark suite. Time to final transcription on the non-streaming endpoint reportedly improves 70% over Chirp 3, the previous model Google shipped for the same workload. That said, the streaming figure is not a regression. Live transcription trades a slice of accuracy for low-latency bidirectional audio, and the 4.0% figure sits well below most production voice stacks. Teams picking a path need to decide which tradeoff fits their product.

Language Coverage and Code-Switching
Automatic language detection covers more than 85 languages, including mid-sentence code-switching. For multilingual contact centers, transcription services, and global voice products, that single capability removes a layer of routing logic that older systems forced developers to build themselves.
Worth noting: the 85-language figure is the detection count, not a separate per-language accuracy table. A handful of low-resource languages are likely to show higher WER than the 2.6% headline average, but Google has not published a per-language breakdown at launch.
Pricing, Access, and the Open-Weights Question
Is it deployable? Yes, but API-only. There are no open weights and no self-hosted path. This is a managed-service decision, not an infrastructure one. Solo developers and startups can start on the Gemini API free tier via Google AI Studio.
Mid-market teams move to the paid tier for higher rate limits, and the paid tier also guarantees content is not used to improve Google’s products. Regulated enterprises route through the Gemini Enterprise Agent Platform, which adds provisioned throughput, compliance controls, and volume discounts.
Both developer and enterprise tracks are reportedly in public availability as of launch.
What This Changes for Voice AI Builders
For voice-agent startups, the practical question is whether Gemini 3.5 Transcribe can replace an existing stack like Whisper, AssemblyAI, or Deepgram without a regression in tail-latency accuracy. The 2.6% non-streaming WER is competitive at the top of the field, and the 85-language detection set is broader than most rivals publish.
The streaming endpoint at 4.0% gives Google a defensible position in the live-agent market without forcing accuracy-sensitive batch jobs to compromise. The next step is independent benchmarking. Artificial Analysis reportedly supplied Google’s headline numbers, and third-party replication across noisy, accented, and code-switched audio will determine whether the 2.6% average holds in production.
Related Articles
- Gemini 3.5 Transcribe: Google’s Most Precise Speech AI Yet
- boAt brings Google Gemini to Its Next-Gen Audio Devices
- Barret Zoph Returns to Google DeepMind After Six Years and Two
FAQs
What is Gemini 3.5 Transcribe?
Gemini 3.5 Transcribe is Google AI’s speech-to-text model released on August 27, 2026, with two endpoints for pre-recorded and live audio transcription.
How accurate is Gemini 3.5 Transcribe?
Google reportedly reports an average WER of 2.6% for non-streaming and 4.0% for streaming transcription, as measured by Artificial Analysis.
How many languages does Gemini 3.5 Transcribe support?
The model supports automatic detection across more than 85 languages, including mid-sentence code-switching.
Can developers self-host Gemini 3.5 Transcribe?
No. The model is reportedly API-only, with no open weights or self-hosted deployment path at launch. Google’s release of Gemini 3.5 Transcribe reportedly sets a new bar for managed speech-to-text — developers should benchmark the 2.6% WER against their own audio before committing to migration.
Was this article helpful?
Your feedback directly improves future articles on this site.





