Google is releasing two voice-first models: Gemini 3.8 Live (optimized for cost and conversational speed) and 3.8 Live Extended Thinking (for reasoning-heavy tasks). Both handle near-real-time dialogue, execute API calls in the background while staying conversational, and support 97 languages mid-stream. The Extended version explicitly reasons aloud to users during complex workflows. Rollout is immediate across Gemini API, Google AI Studio, and consumer products (Search, Workspace). The source cites specific benchmarks—82.6 on Artificial Analysis' Speech-to-Speech index, 97.7% on Big Bench Audio—and names developer integrations (LangChain, Vercel, LiveKit, etc.) already building on the Live API. All generated audio gets SynthID watermarking. For practitioners: if you need to deploy voice agents at scale, the Gemini API gets two new SKUs starting today; the Enterprise tier offers private preview access. One sharp gap: the source doesn't specify latency numbers or token costs, which would typically drive adoption decisions.
reply