Google's Gemini 3.5 Transcribe pushes speech-to-text with sub-second streaming
Google has launched Gemini 3.5 Transcribe, a speech-to-text model it positions as a step beyond conventional recognition engines. Rather than producing raw transcripts, the model cleans up disfluencies on the fly—stripping filler words, resolving self-corrections, and auto-formatting output—while claiming resilience to background noise and specialized jargon. Google reports vendor-measured word error rates of 4.0% for streaming and 2.6% for pre-recorded audio (per Artificial Analysis), plus a 70% improvement in time-to-final-transcription over its prior Chirp 3 model. It handles over 85 languages, supports custom vocabulary, and attributes speech for up to three speakers with word-level timestamps.
The release is aimed squarely at developers, exposed through two APIs: a Live API for real-time bidirectional streaming and an Interactions API for batch processing of recordings and call logs. Integrations with platforms like LiveKit, Pipecat, LangChain, and Vercel are meant to lower the barrier for building voice agents, live captioning, and post-call analytics. A notable wrinkle is function calling—the model can hand off tasks such as image generation or file analysis to other Gemini models, currently in the macOS Gemini app.
Google is also threading the model into consumer surfaces including Gboard’s Rambler feature, Google AI Studio, the macOS Gemini app, and soon Chrome. Worth flagging for a technical audience: the benchmark figures are vendor-supplied, and several features (notably Antigravity’s transcription) depend on the model reading screen context and chat history with user permission—an accuracy-versus-privacy tradeoff developers should weigh. The model is in public preview across Google’s AI Studio, Antigravity, and Enterprise Agent Platform.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.