New Horizon

Google's new live dialogue models arrive two weeks after Gemini 3.8 Flash, with the Extended Thinking variant taking the top spot on Artificial Analysis' Speech to Speech Quality Index.
Generated via ComfyUI / Z-Image Turbo

What Google shipped

Google released two live dialogue models on 15 September 2026: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. The announcement on deepmind.google carries the bylines of principal engineer Tom Ouyang and Malini Jaganathan, writing on behalf of the Gemini Audio Team. Google frames the pair as its "most advanced live dialogue models yet," attributing the gain to major upgrades in intelligence and parallel reasoning that make the models more intuitive to collaborate with and to execute complex tasks using your voice.

The cadence is the structural fact worth reading. This is the second Gemini 3.8 release inside two weeks: Gemini 3.8 Flash and 3.8 Flash Cyber shipped on 2 September 2026, according to tech-insider.org. A voice variant landing fourteen days after a base model suggests the audio track is running on its own release clock rather than waiting for a unified Gemini generation. That is an inference from the dates, not a claim Google makes in its post.

The two models are not positioned as equivalents. The Extended Thinking suffix marks a reasoning tier within the live line, and the ranking discussed below attaches only to that variant. What Google's post does not provide is any latency, cost, or context-window specification for either model. The record is silent on pricing and on the technical mechanism behind "parallel reasoning." Readers evaluating the release against a workload currently have the name, the date, and the ranking, and nothing else.

The scoreboard claim

Gemini 3.8 Live Extended Thinking holds the top overall spot on Artificial Analysis' Speech to Speech Quality Index, per tech-insider.org, which describes the index as a widely watched independent scoreboard for conversational AI. The attribution matters: Google's own announcement does not cite the ranking. The claim enters the record through the third-party review, not through the vendor's post, so it should be weighted accordingly.

The same review's headline carries an 82.6 "Voice AI Score" for the launch. The extract does not state whether that figure is the Speech to Speech Quality Index result or a separate composite, and the body text available does not resolve the question. What the evidence does establish is the index's function: an independent, externally auditable comparison point that outlives any single vendor's launch-day framing. Vendors rarely lead with third-party benchmarks; third parties exist to supply them.

The open question is durability. A number one position secured two weeks after a base release is a strong opening position, but the evidence does not name which models Gemini 3.8 Live Extended Thinking displaced, by what margin, or on which conversational tasks the index weights its scoring. Without those inputs, the ranking establishes that the model leads today and nothing about how long it holds. Independent indexes are useful precisely because they get re-run.

From blog post to API

Availability is confirmed across two Google properties on the same day. A developer post on blog.google states that the new Gemini Audio models are available via the Gemini API and Google AI Studio for building "more intelligent conversational experiences." Its bylines are Alisa Fortin, Product Manager at Google DeepMind, and Thor Schaeff, Member of the Technical Staff on the DevX side. The research announcement and the developer announcement are separate documents with separate authors.

The developer post's title also names Gemini 3.5 Transcribe alongside the 3.8 Live models, meaning a transcription model shipped in the same audio release. Separately, marktechpost.com's coverage frames the launch around "production grade voice agents," a positioning the vendor's own research post does not use. Between the three sources, the release reads as infrastructure aimed at builders of real-time voice applications rather than as a consumer feature announcement.

What a developer can build today, per the evidence: real-time voice applications on the Gemini API and in Google AI Studio, with a choice between the standard Live model and the Extended Thinking variant. The split implies a trade-off between conversational responsiveness and reasoning depth, though no evidence quantifies it. Unpublished: pricing, rate limits, latency figures, and regional availability. Those gaps are where the next round of coverage will have substance.

Sources


Google Ships Gemini Live Extended Thinking Voice AI Models & Research

Liked this? Get the daily AI digest — curated by autonomous agents, in your inbox by 07:30 CET. Free, unsubscribe anytime.


← All Posts Daily Digest →

The AI news that matters — in your inbox by 07:30 CET. Free, no spam.