The future of business, today.
RSSNewslettersAdvertise
Business Future Today

Real-time AI

Google Targets Production Voice Agents With Gemini 3.8 Live

Google’s new Gemini Live models combine real-time voice, visual context, background tool use and—on the higher-end version—multi-step reasoning. The practical question for enterprises is whether those capabilities can improve completion rates without making customer interactions harder to govern.

A chart showing EVA Bench Experience to Task Completion

Google has introduced two real-time dialogue models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, aimed at developers and enterprises building voice-driven AI systems.

The split is deliberate. Gemini 3.8 Live is positioned for scaled, cost-conscious deployments, while Extended Thinking is designed for more complex, multi-step work. Both are available through the Gemini Live API and Google AI Studio; enterprise access is initially in private preview in Gemini Enterprise.

What changed

The release brings several capabilities into one live interaction stack: near-real-time voice dialogue, visual input, tool and API calling, and multilingual conversation. Google says Gemini 3.8 Live can shift automatically among 97 languages within a conversation and can continue talking while tools or API requests run in the background.

That last point matters more than it sounds. Conventional voice bots often force a visible pause while a system fetches an answer, updates a record or completes a workflow. Google’s approach is intended to let the agent acknowledge the request, keep the interaction moving, then return with the result.

Gemini 3.8 Live Extended Thinking adds a different interaction pattern: it can reason through a complex task while speaking, according to Google, offering early acknowledgements and progress narration during multi-step work. The company is positioning that model for tasks where a voice interface must do more than retrieve information—such as coordinating tools and working through a process with the user.

Google cites its own results across several external benchmarks, including an 82.6 score on Artificial Analysis’ Speech-to-Speech Quality Index for Extended Thinking, plus results on τ-Voice, Sierra’s banking benchmark and ServiceNow’s EVA-Bench. Those figures are useful directional signals, but operators should treat benchmark performance as a starting point rather than evidence of performance in their own workflows.

Why it matters to builders

The release reflects a shift in voice AI from scripted automation toward interfaces that must manage conversation, context and action at the same time. For a customer-service, field-support or internal-helpdesk deployment, the important design unit is no longer just the model response. It is the full system: voice latency, tool permissions, escalation paths, data access and auditability.

Google is also reducing integration friction by working through existing real-time AI and media platforms including Agora, LangChain, LiveKit, Pipecat and Vercel. That gives teams options to use established streaming and observability infrastructure rather than building all real-time media plumbing themselves.

For organizations already using Google Workspace, Gemini’s distribution is notable. Google says Extended Thinking is coming to Workspace business customers and is available to eligible subscribers in tools including Docs, Gmail and Keep. Search Live is also receiving Gemini 3.8 Live for real-time troubleshooting assistance. These placements could make voice interaction a more routine entry point to enterprise information and workflows—not a separate bot project.

A practical deployment choice

Teams should separate use cases before picking a model. The standard Live model may fit high-volume conversational tasks where responsiveness and operating cost matter most. Extended Thinking may warrant testing where the business value comes from completing a multi-step workflow correctly, but it also increases the need for careful evaluation of failures, handoffs and tool-call controls.

A useful pilot should measure more than response quality: task completion, time to resolution, transfer-to-human rate, tool-call errors, customer abandonment and the cost per completed task. Test multilingual and interruption-heavy calls, as well as cases where backend systems are slow or unavailable.

What to watch next

The immediate issue is availability: the models are broadly accessible to developers, but enterprise offerings remain partly in private preview or marked “coming soon.” Pricing details were not included in Google’s announcement, despite its claims of cost efficiency, so deployment economics will need scrutiny.

Governance is another test. Google says audio generated by its AI products is watermarked with SynthID, which may help with provenance. But enterprises will still need their own policies for disclosure, recording consent, sensitive data, action authorization and review of consequential outcomes.

The real measure of this release will be whether Google’s live models can sustain natural conversation while reliably finishing useful work in production—not simply sounding more human while they try.

Sources

STAY AHEAD

The future of business, in your inbox.

Useful signals on the companies, technologies and shifts changing business.

One useful briefing. Unsubscribe any time.