Sep 16 edition/Reporting & analysis
ModelsAgentsInfrastructureBusinessSafety

ModelsArchitectures & capability

Google splits Gemini 3.8 Live into fast voice and Extended Thinking modes for real-time agents

Google’s Gemini 3.8 Live launch divides real-time speech agents into a low-latency interaction model and an Extended Thinking mode for longer, tool-using workflows. The key technical shift is a more complex session lifecycle, with benchmarks and prior research pointing to the need for task-specific validation.

Illustration from Simon Willison’s Weblog: Google splits Gemini 3.8 Live into fast voice and Extended Thinking modes for real-time agents
Image: Simon Willison’s Weblog — Original article ↗
THE CORE IDEAS4 TAKEAWAYS
01

Google says Gemini 3.8 Live is aimed at fast speech-to-speech interaction, while Gemini 3.8 Live Extended Thinking is designed for longer spoken tasks involving reasoning and tool use. [7] [14]

02

Extended Thinking changes client design: developers cannot treat a completed utterance as the end of the whole request and must track interaction state and non-blocking tool execution. [14]

03

The Live API is exposed over WebSockets, and Simon Willison’s prototype suggests browser-based testing is straightforward, but production systems still need secure auth, interruption handling, observability, and media infrastructure. [1] [8]

04

Independent benchmark listings are useful but limited, and both the model card and prior voice-AI research point to unresolved risks around hallucination, latency, tool reliability, and sensitivity to vocal cues. [2] [4] [5] [13]

WHY IT MATTERS

Evidence in the reviewed research shows a real developer-facing shift: Google’s docs define separate model modes, lifecycle states, WebSocket access, pricing, and known limitations, while Willison’s demo adds narrow practitioner evidence that a minimal client can work.

Read the full assessment

The implication for businesses is not just better voice UX; it is the possibility of agents that keep speaking while tools run. That also raises procurement and engineering demands for logging, cancellation, safety review, and workflow-specific evaluation.

Executive brief

On September 15, 2026, Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two real-time speech-to-speech Gemini models aimed at low-latency voice agents and more complex spoken workflows. Google’s developer documentation says Extended Thinking changes client lifecycle handling: turnComplete no longer means the whole user request is done; developers must track interactionStatus values such as IN_PROGRESS and IDLE. Thinking in the Live API | Gemini API | Google AI for Developers Google’s claims about language switching, visual input, Workspace/Search rollout, SynthID watermarking, and availability are vendor-reported.

Read the full section

On September 15, 2026, Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two real-time speech-to-speech Gemini models aimed at low-latency voice agents and more complex spoken workflows. Simon Willison’s same-day commentary is useful because it adds practitioner evidence: he built and published a minimal browser UI that connects directly to Google’s Live API WebSocket endpoint, captures and plays audio with the Web Audio API, and supports interruption while the model is speaking. Tool: Gemini Live audio

The core product change is not simply “better voice.” Google is positioning two distinct operating modes: gemini-3.8-live for fast, fluid, low-latency interactions, and gemini-3.8-live-extended-thinking for longer-running, multi-step, tool-using voice tasks where the agent can speak progress/filler updates while reasoning and calling tools in the background. Google’s developer documentation says Extended Thinking changes client lifecycle handling: turnComplete no longer means the whole user request is done; developers must track interactionStatus values such as IN_PROGRESS and IDLE. Thinking in the Live API | Gemini API | Google AI for Developers

Evidence is mixed in strength. Google’s claims about language switching, visual input, Workspace/Search rollout, SynthID watermarking, and availability are vendor-reported. Gemini 3.8 Live & Gemini 3.8 Live Extended Thinking Artificial Analysis provides an independent benchmark source for its Speech-to-Speech Index methodology and current leaderboard, but its aggregate score should be read as a benchmark result, not proof of production reliability in a specific enterprise workflow. Speech to Speech Models and Providers Analysis | Artificial Analysis Independent academic work on earlier real-time voice systems warns that current voice AIs may still fail to use prosody, emotion, and delivery cues even when they process audio directly; that finding was not about Gemini 3.8 specifically, but it is a relevant caution for high-stakes voice-agent deployment. Real-Time Voice AI Hears but Does Not Listen

What changed and event timeline

  1. Google published its launch post for Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking

    The post describes Gemini 3.8 Live as optimized for scale, cost efficiency, conversational intelligence, fluid dialogue, and visual grounding, while Extended Thinking is described as intended for higher-complexity tasks with multi-step reasoning.

  2. Simon Willison posted “Gemini Live audio,” classifying it as a tool/demo note

    He reports that Google released the two models that day and that he had a model generate a library-free web UI for testing them.

    More detail

    His implementation evidence is narrow but useful: it confirms that a browser client can use the documented WebSocket API directly and that interruption behavior can be exercised from a simple prototype.

  3. Current access

    Google says both models are rolling out through the Gemini API and Google AI Studio for developers.

    More detail

    Google also says Gemini 3.8 Live is rolling out in Search Live, while Extended Thinking is rolling out in Gemini Live and selected Workspace experiences, with enterprise access in private preview and broader enterprise/customer-experience availability “coming soon.”

Capabilities and access

Known model IDs are: The Google DeepMind model card says the Gemini 3.8 Audio models are based on Gemini 3 Pro, accept audio/images/video/text with up to a 128K-token context window, and output audio and text with a 64K-token output limit. The card does not disclose enough architecture or training detail to independently reproduce the model.

Read the full section

Known model IDs are:

  • gemini-3.8-live — stable model string, latest update September 2026. Google’s model page lists text, image, audio, and video inputs; text and audio outputs; 131,072 input tokens; 65,536 output tokens; Live API support; function calling; search grounding; audio generation; and no support for caching, code execution, file search, image generation, structured outputs, or URL context. Gemini 3.8 Live | Gemini API | Google AI for Developers
  • gemini-3.8-live-extended-thinking — Google’s Live API “Thinking” docs identify this as the model for background reasoning in real-time voice sessions, with configurable thinking levels of low, medium, and high; MINIMAL is not supported. Thinking in the Live API | Gemini API | Google AI for Developers

The Google DeepMind model card says the Gemini 3.8 Audio models are based on Gemini 3 Pro, accept audio/images/video/text with up to a 128K-token context window, and output audio and text with a 64K-token output limit. The card does not disclose enough architecture or training detail to independently reproduce the model. Gemini 3.8 Audio (Live, Live Extended Thinking) - Model Card — Google DeepMind

Pricing on Google’s developer pricing page groups Gemini 3.8 Live, Gemini 3.8 Live Extended Thinking, and Gemini 3.1 Flash Live Preview together. For paid standard API use, Google lists audio input at $3.00 per 1M tokens or $0.005/min, audio output at $12.00 per 1M tokens or $0.018/min, text input at $0.75 per 1M tokens, and text output at $4.50 per 1M tokens. The same table says free-tier content is used to improve Google products, while paid-tier content is not. Gemini Developer API pricing | Gemini API | Google AI for Developers

Technical analysis for researchers and developers

The most important developer-facing change is the move from a simple “turn is done” interpretation to a richer session lifecycle for Extended Thinking. For standard Gemini 3.8 Live, Google says turnComplete: true closes the model’s turn and returns the session to idle.

Read the full section

The most important developer-facing change is the move from a simple “turn is done” interpretation to a richer session lifecycle for Extended Thinking. For standard Gemini 3.8 Live, Google says turnComplete: true closes the model’s turn and returns the session to idle. For Extended Thinking, Google says the model can emit spoken fillers, tool calls, and final answers across one user request; turnComplete: true may only mean an utterance is finished, while interactionStatus determines whether the task is still in progress. Thinking in the Live API | Gemini API | Google AI for Developers

This matters for UI state machines. A naive client that resumes listening, enables form submission, or tears down spinners when it sees turnComplete may behave incorrectly with Extended Thinking. Production clients should instead treat interactionStatus: "IN_PROGRESS" as a continuing task state and interactionStatus: "IDLE" as the point at which the full multi-step interaction has completed. Thinking in the Live API | Gemini API | Google AI for Developers

Tool integration also changes. Google says Extended Thinking requires function declarations with behavior: "NON_BLOCKING"; synchronous blocking tools return an error. That shifts responsibility to the application layer: developers need correlation IDs, async tool-response handling, cancellation semantics, timeout management, user-visible progress, and guardrails for tools that may execute after the user has interrupted or changed intent. Thinking in the Live API | Gemini API | Google AI for Developers

The raw WebSocket path is documented and relatively approachable. Google’s getting-started guide says the Live API supports real-time bidirectional interaction with audio, video, and text inputs and native audio outputs over WebSockets; the first message must be a BidiGenerateContentSetup, and communication uses JSON messages conforming to client/server message structures. Get started with Gemini Live API using WebSockets | Gemini API | Google AI for Developers Simon Willison’s demo reinforces that a browser-only implementation is feasible without third-party libraries, though that should not be confused with a production architecture: exposing API keys in browser clients is usually inappropriate, and Google separately documents ephemeral authentication tokens for constrained Live API sessions. Tool: Gemini Live audio

For audio transport, Google’s Thinking guide documents 16 kHz raw PCM for real-time input chunks and 24 kHz PCM for streamed model audio output in the examples. Thinking in the Live API | Gemini API | Google AI for Developers That has implementation implications: browser apps need sample-rate conversion, buffering, playback scheduling, echo management, and graceful interruption. Willison’s Web Audio API prototype is a useful minimal reference, but enterprise deployments will likely use WebRTC/media infrastructure, server-side relays, or vendor platforms for NAT traversal, observability, moderation hooks, and session recording policies. Tool: Gemini Live audio

Claims and evidence

  • Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026.
  • gemini-3.8-live is the stable model string for the low-latency Live model. — Vendor documentation.
  • Extended Thinking supports background reasoning, spoken fillers, and async tools. — Vendor documentation.
Read the full section
Material claimEvidence status
Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026.Vendor-reported by Google; independently noted by Willison the same day. Gemini 3.8 Live & Gemini 3.8 Live Extended Thinking
gemini-3.8-live is the stable model string for the low-latency Live model.Vendor documentation. Gemini 3.8 Live | Gemini API | Google AI for Developers
Extended Thinking supports background reasoning, spoken fillers, and async tools.Vendor documentation. Thinking in the Live API | Gemini API | Google AI for Developers
Extended Thinking scored 82.6 on Artificial Analysis’ Speech-to-Speech Index and 68.6% on τ-Voice in the current AA table.Independent benchmark publication by Artificial Analysis, also cited by Google. Not a production guarantee. Speech to Speech Models and Providers Analysis | Artificial Analysis
Artificial Analysis’ index is an equal-weighted composite across speech reasoning, agentic performance, arena preference, and task success.Independent methodology page. Speech Reasoning Benchmarking Methodology | Artificial Analysis
Gemini 3.8 Audio may hallucinate and may experience occasional slowness or timeouts.Vendor model card limitation. Gemini 3.8 Audio (Live, Live Extended Thinking) - Model Card — Google DeepMind

Context and prior work

Gemini 3.8 Live arrives in a competitive cycle focused on native audio, full-duplex, and delegated/background reasoning. Benchmarking is also evolving. Full-Duplex-Bench-v3 similarly focuses on naturalistic speech disfluency and multi-step tool use, evaluating older systems including Gemini Live 2.5 and Gemini Live 3.1 rather than Gemini 3.8.

Read the full section

Gemini 3.8 Live arrives in a competitive cycle focused on native audio, full-duplex, and delegated/background reasoning. OpenAI’s July 2026 GPT-Live launch framed earlier voice systems as cascades of speech-to-text, LLM, and text-to-speech, and described GPT-Live as a full-duplex architecture that can listen and speak continuously while delegating deeper work to frontier models. Introducing GPT-Live | OpenAI OpenAI later described a production architecture that separates the real-time media path from slower application/tool logic, an approach conceptually similar to Google’s Extended Thinking distinction between spoken interaction and background work, though the implementation details and model architectures differ and should not be assumed equivalent. How we built a realtime system for responsive voice AI in six months | OpenAI

Benchmarking is also evolving. Sierra’s τ-Voice paper evaluates voice agents using customer-service scenarios with interruptions, frame drops, background noise, muffling, non-agent-directed speech, and other realistic events; it measures task success plus responsiveness, latency, interrupt rate, and selectivity. $\tau$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains Full-Duplex-Bench-v3 similarly focuses on naturalistic speech disfluency and multi-step tool use, evaluating older systems including Gemini Live 2.5 and Gemini Live 3.1 rather than Gemini 3.8. Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency

Limitations, safety, and contested findings

The model card’s limitations are broad but important: hallucinations remain possible, jailbreak resistance is still an active work area, occasional slowness/timeouts may occur, and the stated knowledge cutoff is January 2025. Gemini 3.8 Audio (Live, Live Extended Thinking) - Model Card — Google DeepMind Google says all audio generated by its AI products is watermarked with SynthID.

Read the full section

The model card’s limitations are broad but important: hallucinations remain possible, jailbreak resistance is still an active work area, occasional slowness/timeouts may occur, and the stated knowledge cutoff is January 2025. Gemini 3.8 Audio (Live, Live Extended Thinking) - Model Card — Google DeepMind Google also says Gemini 3.8 Audio went through automated, human, and red-team-style safety evaluations, but the public model card does not provide enough detail to independently audit those results. Gemini 3.8 Audio (Live, Live Extended Thinking) - Model Card — Google DeepMind

Google says all audio generated by its AI products is watermarked with SynthID. That is a useful provenance feature, but the launch post’s statement should be treated as a vendor claim unless tested under real distribution, transcoding, recording, and adversarial-edit conditions. Gemini 3.8 Live & Gemini 3.8 Live Extended Thinking

A key contested area is whether native audio models truly “listen” beyond transcripts. A 2026 paper evaluating production real-time systems, including an older Google Gemini 3.1 Flash Live system, found that systems often acted on words rather than vocal delivery in scenarios where tone and delivery carried meaning. This does not establish that Gemini 3.8 fails in the same way, but it argues for domain-specific evaluation before using live voice agents in mental health, fraud, negotiation, accessibility, education, or safety-critical support. Real-Time Voice AI Hears but Does Not Listen

Business and practitioner implications

For business leaders, the practical opportunity is voice agents that can keep the user engaged while doing real work: checking records, calling tools, searching, drafting, or coordinating workflows. For developers, Gemini 3.8 Live looks attractive for prototypes and low-friction integration because the API is exposed over WebSockets and documented with raw JavaScript/Python examples.

Read the full section

For business leaders, the practical opportunity is voice agents that can keep the user engaged while doing real work: checking records, calling tools, searching, drafting, or coordinating workflows. The risk is that “speaks while thinking” can mask uncertainty, latency, or tool failure behind confident conversational progress. Procurement teams should require logs that distinguish model speech, tool calls, tool results, safety interventions, and user interruptions.

For developers, Gemini 3.8 Live looks attractive for prototypes and low-friction integration because the API is exposed over WebSockets and documented with raw JavaScript/Python examples. Get started with Gemini Live API using WebSockets | Gemini API | Google AI for Developers For production, the main engineering burden is not making audio play; it is building a reliable state machine, identity/authentication boundary, observability, transcript policy, interruption/cancellation behavior, and async tool governance.

For researchers, Gemini 3.8 is a useful target for evaluation because it exposes the central frontier problem in voice agents: balancing latency, turn-taking, tool use, reasoning quality, and truthful progress reporting. Artificial Analysis’ current numbers are encouraging, but they should be supplemented with reproducible, task-specific tests using real user accents, noisy channels, domain policies, and failure-mode annotation. Speech to Speech Models and Providers Analysis | Artificial Analysis

Sources

Read the full section
FOLLOW THE EVIDENCE

The source trail.

Sources (15)
A LITTLE LESS NOISE. A LOT MORE CONTEXT.

Stay curious.
Follow the evidence.

Independent perspectives, the original sources, and room for the questions that don't have easy answers.

How we build the brief