Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages
MarkTechPost
Read full postGoogle has launched Gemini 3.5 Transcribe, a speech-to-text model supporting over 85 languages with a 2.6% average word error rate for recorded audio and 4.0% for streaming. It offers two APIs: one for pre-recorded files and another for live streaming, each with distinct features and limits. The model is accessible via API only, with no open weights or self-hosting options, targeting industries like contact centers, clinical documentation, and media captioning.

.png?disable=upscale&width=1200&height=630&fit=crop)

