RC RANDOM CHAOS

Google's Gemini 3.8 Live models push for production-grade voice agents

· via Hacker News

Original source

Gemini 3.8 Live and 3.8 Live Extended Thinking

Hacker News →

Google has released two speech-to-speech models aimed at making voice agents viable in production. Gemini 3.8 Live is the cost-efficient tier built for scale, handling fluid dialogue, near real-time visual input, and automatic switching across 97 languages mid-conversation. Its sibling, 3.8 Live Extended Thinking, targets harder multi-step work by reasoning and speaking at the same time — filling dead air with cues like ‘Let me check that…’ and narrating progress while tool calls run in the background.

Google leans heavily on benchmark placement to make its case: Extended Thinking claims the top spot on Artificial Analysis’ Speech-to-Speech Quality Index (82.6) and leads agentic task-completion tests like τ-Voice and Sierra’s banking variant, while the lighter model ranks second in the Speech Agent Arena. The pitch is that both models balance conversational quality against accuracy and price, with the Live API doing the tool execution and background task handling that separates a demo from a deployable agent.

Distribution is the real story. The models plug into voice-infrastructure platforms like LiveKit, Pipecat, LangChain, and Vercel, and land inside Gemini, Workspace, and Search for end users, with enterprise access in private preview. All generated audio carries Google’s SynthID watermark, an imperceptible signal meant to keep AI-generated speech detectable — a nod to the misinformation risk that comes with making synthetic voices this fluent.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.