Gemini 3.8 Live: Google Undercuts OpenAI Sharply on Price

Sprachassistent-Gerät auf einem Schreibtisch als Sinnbild für KI-gestützte Sprachagenten
Photo by BENCE BOROS on Unsplash

Voice assistants that can handle background tasks mid-conversation without losing the thread have mostly been an OpenAI talking point so far. Google DeepMind is now catching up with two new models — and positioning them explicitly as faster and significantly cheaper than the competition. For developers building voice agents, that’s more than a technical footnote: the price gap with OpenAI’s offering is large enough to decide whether entire products are economically viable.

Key takeaways

  • Google DeepMind is releasing two new audio models for voice agents, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, that can call tools in the background while continuing to talk.
  • Gemini 3.8 Live Extended Thinking tops Artificial Analysis’s Speech-to-Speech leaderboard with a score of 82.6, narrowly ahead of OpenAI’s GPT-Live-1 Astra.
  • An hour of voice conversation on Google’s standard model costs around $0.84 — roughly a sixth of what OpenAI’s comparable offering charges.
  • Both models support more than 97 languages, process visual input in near real time, and watermark generated audio output with SynthID.
  • Early partners including Salesforce, ServiceNow, and Lenskart are already using the models in production; they’re available through the Gemini API, Google AI Studio, and gradually across Google’s own products.

Two models, one shared core

Gemini 3.8 Live is aimed at applications where cost and scalability matter most, according to Google: “Built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding,” the Gemini Audio Team, led by principal engineer Tom Ouyang, writes in the official blog post. The second variant, Gemini 3.8 Live Extended Thinking, targets more complex tasks and includes an extra reasoning step that enables multi-step reasoning — useful for, say, nested booking flows or coordinating several asynchronous function calls. Unlike a purely sequential design, that reasoning step isn’t meant to create noticeable conversational pauses: the model narrates its intermediate steps out loud, with phrases like “Let me check that…”, while it keeps working in the background.

What the models can do mid-conversation

The central advance is that both models can execute tools and API calls during an ongoing conversation without interrupting it. A user can make a request while the system works on a task in the background, confirms the result, and keeps talking normally. On top of that comes near-real-time processing of visual input: in demos, the model walks new employees through onboarding while taking into account what’s currently on screen, or turns a hand-drawn sketch plus spoken feedback into a working React component. Language support automatically covers more than 97 languages, including mid-conversation switching. Google gives no concrete figures for latency in milliseconds or context window size — unlike OpenAI’s GPT-Live-1, which according to its own announcement relies on true full-duplex operation, letting users and the model speak and interrupt each other simultaneously. Google describes Gemini 3.8 Live Extended Thinking only as being able to “reason and speak simultaneously” and handle interruptions, without using the term full-duplex itself.

Price and performance compared

On Artificial Analysis’s benchmarks, Gemini 3.8 Live Extended Thinking leads the Speech-to-Speech quality index with 82.6 percent, narrowly ahead of OpenAI’s GPT-Live-1 Astra at 81.5 percent and xAI’s Grok Voice Think Fast 2.0 at 81.3 percent. On agentic tasks, measured by the τ-Voice benchmark, Google scores 68.6 percent versus 67.9 percent for OpenAI and 56.5 percent for xAI; on Sierra’s specialized τ-Voice-banking test, Google also leads with 35.1 percent. The real difference shows up in price, though: an hour of voice conversation costs around $0.84 on Gemini 3.8 Live’s standard mode, and about $3.50 on the Extended Thinking variant. OpenAI’s GPT-Live-1 Astra, by contrast, runs about $5.83 an hour, while xAI’s Grok Voice Think Fast 2.0 costs $4.80. Google’s standard model thus costs roughly a sixth of what OpenAI’s comparable offering charges — while claiming higher benchmark quality. How these numbers hold up under real users and variable conversation lengths remains to be seen; Artificial Analysis figures reflect lab tests, not production load.

Who’s already using the models

In its blog post, Google names a number of partners already using or testing the new models in production, including Salesforce, ServiceNow, Indian e-commerce platform Lenskart, and healthcare provider Lumeris. Several of these customers, according to Google, praise the latency, conversational fluidity, and reliable tool-calling in particular — though Google doesn’t share concrete figures, and the statements remain qualitative. The models are available to developers immediately through the Gemini API and Google AI Studio, and to enterprise customers initially as a private preview on the Gemini Enterprise platform. End users will encounter the standard variant already in Google’s Search Live, while the Extended Thinking version is set to roll out gradually to Gemini Live, Workspace services like Docs for paying subscribers, and Gmail and Keep. Developers who want to experiment can find sample apps in Google’s official GitHub repository.

What this means for developers

For teams already building voice agents — in customer service, appointment booking, or internal support tools — the price gap noticeably changes the math: applications that were economically questionable at OpenAI’s rates, like sustained conversations with many simultaneous users, could suddenly become profitable at Google’s prices. As with OpenAI’s earlier push with GPT-Live-1, aimed mainly at replacing traditional hold music, this shows that competition for voice agents is increasingly being decided by cost per conversation minute, not benchmark scores alone. But anyone who needs especially natural, interruptible conversations with true full-duplex operation will, based on the current state of things, likely still end up with OpenAI — and knowingly pay the higher price for it.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top