Google released, on 15 September, two audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. The first is pitched at scale and cost efficiency, the second at multi-step reasoning on harder tasks. They are rolling out through the Gemini API and AI Studio for developers, in private preview in Gemini Enterprise, in Search Live for the public, and in Google Workspace for business customers in the Extended Thinking variant. Google has not published pricing for either.
The stated numbers: 82.6 on Google’s Speech to Speech Quality Index for Extended Thinking, 68.6% agentic task completion on τ-Voice and 35.1% on Sierra’s τ-Voice-banking, and 97.7% on Big Bench Audio reasoning. Gemini 3.8 Live is placed second on Speech Agent Arena. Ninety-seven languages are supported, with switching detected mid-conversation, and visual input is processed in near real time.
The banking figure is the one to read. A voice agent that finishes 35.1% of banking tasks is a long way from one a bank could put in front of customers unsupervised, and Google’s own launch post says so — which is more useful than the 97.7% audio-reasoning score. τ-Voice-banking is Sierra’s benchmark rather than Google’s, while the Speech to Speech Quality Index is Google’s own measure, so the two are not the same kind of claim.
