Ai Technology News Global

Google Launches Gemini 3.8 Live and Extended Thinking: A New Era for Voice AI Agents

Google has introduced its most advanced live dialogue models yet Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking designed to make voice conversations with AI more natural, intelligent, and capable of handling complex, multi-step tasks in real time.

A modern infographic poster titled "Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking," showcasing vibrant teal UI mockups of a woman using voice assistance
Google has officially launched its advanced Gemini 3.8 Live and 3.8 Live Extended Thinking models, bringing native speech-to-speech interaction, natural dialogue, and real-time background reasoning to developers.

Executive summary

On September 15, 2026, Google rolled out two new native speech-to-speech models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, a pair of live dialogue models rolling out across the Gemini API, Google AI Studio, Gemini Enterprise, Search Live, Gemini Live, and Google Workspace.

The models were built specifically for voice-first interactions, allowing users to hold natural spoken conversations while the AI reasons, calls external tools, and processes visual context in the background — all without breaking conversational flow. Both are hosted models rather than open weights, meaning there is no self-hosted option, though they are live today for production use via the Gemini Live API and Google AI Studio.

What Are Gemini 3.8 Live and Extended Thinking?

The models were introduced by Tom Ouyang, a principal engineer, and Malini Jaganathan, a member of technical staff, writing on behalf of the Gemini Audio Team, who describe the pair as Google's most advanced live dialogue models yet, built for natural conversation. The company positions Gemini 3.8 Live for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding, while Gemini 3.8 Live Extended Thinking is built for high-complexity tasks that call for increased intelligence and multi-step reasoning.

Unlike traditional voice assistants that rely on separate pipelines for speech recognition, language processing, and text-to-speech conversion, Google positions both models as a streamlined alternative to cascaded speech pipelines that chain ASR, an LLM, and TTS. This native speech-to-speech architecture is designed to reduce latency and make conversations feel more fluid and human-like.

Standout Features

A core differentiator for the Extended Thinking model is its ability to reason without pausing the conversation. It supports configurable thinking to help handle complex, multi-step reasoning in the background, while responding or narrating its progress in the main conversation. Developers integrating the model need to adapt their systems accordingly, since Gemini 3.8 Live Extended Thinking introduces background reasoning during live audio sessions, meaning turnComplete: true no longer indicates that the model is idle.

Multilingual and visual capabilities are also central to the release. The model's real-time language switching is aimed at global enterprise use cases — think customer service agents fielding calls that shift between English, Spanish, and other languages without warning.

Benchmark Performance

Google backed the launch with a string of benchmark claims. Gemini 3.8 Live Extended Thinking leads agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra's τ-Voice-banking benchmark, and also scores 97.7% on Big Bench Audio, a reasoning benchmark for audio models. Meanwhile, Gemini 3.8 Live secured second place in the Speech Agent Arena, a human preference evaluation, and on ServiceNow's EVA-Bench, Google reports that the models push the Pareto Frontier for complex workflows.

Availability and Pricing

Both models are rolling out immediately. They are live today in the Gemini Live API and Google AI Studio, available for API-based production use, though as hosted models there is no self-hosted deployment option. On pricing, Google says the models are competitively priced to allow developers to scale voice applications with industry-leading performance.

Why It Matters

This launch signals Google's continued push toward voice as a primary AI interface rather than a secondary feature bolted onto text-based chat. By enabling background reasoning and tool execution without interrupting dialogue, Google is targeting a real gap in current voice-agent technology: the awkward pause while an assistant "thinks." For businesses building customer-facing voice agents in banking, retail, or enterprise support the combination of low latency, multilingual switching, and asynchronous tool use could meaningfully change how conversational AI is deployed at scale.

References

  1. MarkTechPost: Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking https://www.marktechpost.com/2026/09/15/google-releases-gemini-3-8-live-and-3-8-live-extended-thinking-for-production-grade-voice-agents/
  2. Google DeepMind: Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/

Source for the development reported here: lab-announcements

Cite this

Administrator (2026, September 15). Google Launches Gemini 3.8 Live and Extended Thinking: A New Era for Voice AI Agents. AI News Report. https://ainewsreport.org/blog/google-gemini-3-8-live-extended-thinking-voice-ai