Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking generally available (GA)
By Steven Van ·
Google's Live API gets two GA audio-to-audio models: a low-latency default and an extended-thinking option for heavier reasoning.
Two audio-to-audio models for the Live API are now generally available in Gemini. Both are built for real-time voice applications rather than text chat.
- Gemini 3.8 Live (gemini-3.8-live) is the default model for low-latency voice agents and real-time dialogue without reasoning delays. It uses interleaved reasoning, asynchronous function calling by default, and full session client content updates.
- Gemini 3.8 Live Extended Thinking (gemini-3.8-live-extended-thinking) supports background reasoning during live audio interactions, for use cases that need higher reasoning quality than the default model provides.
Details on architecture and setup are in Google's Live API thinking guide.
Google Gemini
Google's AI assistant — write, plan, brainstorm, generate images, and analyze files with one of the most powerful multimodal models.
View Google Gemini →