Key takeaways

  • Google Cloud declared Gemini 3.8 Live with Live Avatar generally available on September 25, 2026, one day after the product announcement. It is available through Gemini Enterprise and the Live API under model ID gemini-3.8-live.
  • The documented non-global API rates per one million tokens are $0.75 text input, $1 image/video input, $3 audio input, $4.50 text output, $12 audio output and $1 avatar-video output. Google lists no cached-input rate for this model.
  • The model card lists a 128K input context, 64K output for Live, and 24K output when Live Avatar is used. It also says continuous avatar interaction supports a few minutes—not extended hours.
  • Google says preset avatars are available, while custom avatars require enterprise allowlisting and verification. Generated audio and video carry SynthID according to the launch materials.
  • Treat “97 languages” in launch copy and “24 supported languages” in the current API overview as different, unresolved scopes. Verify the exact language, voice, avatar and endpoint combination before committing to a rollout.
01

Google moved Live Avatar into general availability on September 25

Google announced Gemini 3.8 Live with Live Avatar on September 24, 2026. On September 25, 2026, Google Cloud said the feature was generally available in Gemini Enterprise and “ready for enterprise production.” The callable model is gemini-3.8-live. Google documents US and EU endpoints and says provisioned throughput is available. Gemini 3.8 Live Extended Thinking is a different product and remains in private preview.

This is a model and API launch, not merely a demo. Developers can build web, mobile and kiosk experiences that accept live audio, images or video and return spoken responses with a synchronized video avatar. Availability does not establish workload fitness: AccessAllGPT did not authenticate an account or verify entitlement, regional routing, quota or production behavior.

02

The API has six modality-specific prices

Google’s Agent Platform pricing table lists Gemini 3.8 Live API in a non-global region at $0.75 per million text-input tokens, $1 per million image/video-input tokens and $3 per million audio-input tokens. Output is $4.50 per million text tokens, including response and reasoning, $12 per million audio tokens and $1 per million avatar-video tokens. The table displays the same rate above and below 200K input tokens, although the model card separately caps Live context at 128K.

The table marks cached input as unavailable and says input and output are charged only for requests returning HTTP 200. Those are token rates, not a call, minute or completed-conversation price. A useful budget must meter each modality and include retries, tools, idle connection policy, provisioned capacity and the application around the model. AccessAllGPT did not receive a billing receipt and does not infer how many video or audio tokens a representative minute consumes.

03

Live Avatar changes the output contract and the limit

The Gemini 3.8 Audio model card says standard Gemini 3.8 Live accepts audio, images, video and text with a token context window up to 128K, and returns audio and text with up to 64K output. With Live Avatar, output becomes audio, video and text, and the listed output limit falls to 24K. The API overview describes avatar video as MP4 and the session transport as a stateful WebSocket connection.

The same model card says Live Avatar can support a few minutes of continuous interaction rather than extended hours. That sentence is a material product boundary for reception, coaching, claims, education and kiosk plans. Design a visible handoff, restart or non-avatar continuation path; do not promise an open-ended video conversation from a launch reel.

04

The avatar can keep talking while a tool runs

Google says asynchronous function calling lets the model trigger tools or API calls in the background while dialogue continues. The Cloud launch gives a hotel check-in style interaction as an example, and says interruption recovery can preserve conversation context and backend transactions. These are vendor descriptions, not results independently reproduced by AccessAllGPT.

Concurrency creates an application-state problem. The avatar may acknowledge work before a tool succeeds, fail after an optimistic statement, or receive a user correction while the previous action is still running. Keep the tool state authoritative outside the model. Surface pending, succeeded, failed and cancelled states; use idempotency keys; and require a fresh policy check before money, messages, records or permissions change.

05

“97 languages” and “24 supported languages” need scope clarification

The launch article says Live Avatar can transition across 97 languages while adapting lip-sync and expressions, and the Cloud announcement says Gemini 3.8 Live understands and speaks 97 languages with automatic detection. The Live API overview retrieved September 27 instead says “Converse in 24 supported languages.” Google may be describing different layers or levels of support, but the reviewed pages do not reconcile the numbers.

Do not turn the larger launch claim into a support guarantee. For each required language, verify input understanding, generated speech, voice availability, code switching, transcript quality, lip synchronization, policy behavior and fallback. Record whether the result uses a documented supported language, a demonstrated capability or an unsupported observation. The discrepancy itself should remain open until Google publishes a scope map.

06

Custom avatars are not self-service general availability

Google offers a curated library of preset avatars. Creating a custom avatar from a reference image—and, in the Cloud demo, an audio sample—requires enterprise allowlisting and verification. The general availability of the model therefore does not mean every customer can immediately upload a person or brand character and deploy it.

Procurement should separate base-model access, Live Avatar entitlement and custom-avatar approval. Before submitting a likeness, establish documented consent, usage scope, revocation, retention, deletion and incident ownership. Test look-alike rejection and account offboarding. These are AccessAllGPT controls, not claims that Google’s verification fails.

07

SynthID is transparency metadata, not the whole misuse control

Google says all generated Live Avatar audio and video streams carry imperceptible SynthID watermarks. That can support provenance and detection, but the reviewed launch pages do not publish a false-positive rate, false-negative rate, survival result for every transformation or service-level commitment for downstream detection. AccessAllGPT did not generate or inspect a stream.

Make the experience visibly disclose that the participant is AI, even when watermark detection is unavailable. Preserve consent and generation records, restrict export where the use case does not need it, and give users a route to a human. A hidden watermark should complement—not replace—clear presentation, authorization and abuse response.

08

The model card names limits the demos do not show

Google says Gemini 3.8 Audio is based on Gemini 3 Pro and delegates architecture, training-data, implementation and several safety details to that model’s card. It names hallucinations, occasional slowness or timeouts and a January 2025 knowledge cutoff. The “few minutes” avatar-duration limit is also in the card rather than the launch headline.

The card says Google performed automated, human, child-safety and red-team evaluations, and concludes the 3.8 Audio variants are unlikely to reach a tracked or critical Frontier Safety capability level by reference to Gemini 3.7 Flash. This is a vendor safety assessment, not an independent test of a particular avatar application, identity workflow or tool policy.

09

Run a conversation-and-action acceptance test

Start with one preset avatar, one endpoint and a narrow read-only workflow. Freeze cases for interruption, cross-talk, silence, camera loss, language switches, unsafe input, tool delay, tool denial, duplicate tool completion, user correction and the published session-duration boundary. Record first audio, first video, end-to-end turn time, transcript differences, tool-state accuracy, reconnect behavior, tokens by modality and cost per accepted conversation.

Then test the consequence boundary with inert tools. The avatar must not claim completion before durable tool success, repeat a side effect after reconnect, or conceal an incomplete action behind fluent conversation. Evaluate accessibility without assuming a human-like face helps every user: include captions, keyboard operation, reduced-motion needs, screen-reader flow, audio-only fallback and a non-avatar route.

10

The decision: pilot a bounded interaction, not a synthetic employee

Pilot Live Avatar when visual presence materially improves one short, measurable interaction and the team can obtain the needed entitlement, meet consent requirements, meter all modalities and provide a deterministic fallback. Use a preset avatar first. Add a custom avatar only after allowlisting, identity rights and removal procedures are confirmed. Keep consequential tools behind application-owned state and authorization.

Wait when the workflow needs hours-long continuity, an unverified language, guaranteed global routing, a custom likeness without approval, or a cost model that assumes audio and video tokens behave like text. Reject the avatar layer when plain voice, text or a human channel meets the outcome with less identity, accessibility and operational risk. General availability answers whether Google sells the service now; it does not answer whether the service belongs in a particular customer interaction.

11

Copy-ready Live Avatar launch decision record

Complete this record for one interaction, endpoint, language set and avatar. Replace Google claims with account and workload evidence before production approval.

Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].

Pilot, deploy, constrain, wait or reject; user; task; expected duration; consequence; owner; review and expiry dates.

Project, model ID gemini-3.8-live, Gemini Enterprise entitlement, SDK or WebSocket client, US/EU endpoint, quota and provisioned-throughput status.

Preset or custom; allowlist result; likeness and voice rights; approved contexts; disclosure; revocation, deletion and incident owner.

Permitted text, audio, image and video inputs; required audio, text and avatar outputs; 128K context, 24K avatar-output and session-duration handling.

Required languages and code switches; documented support versus launch claim; comprehension, speech, transcript, lip-sync, policy and fallback test results.

Tools, identities, permissions, pending-state presentation, authorization, idempotency, cancellation, timeout, reconciliation and prohibited actions.

Tokens and posted rate for every input/output modality; retries, tools, throughput, application cost; cost per accepted conversation; budget and alert.

Frozen cases; interruption and reconnect results; latency percentiles; captions; keyboard, screen-reader, reduced-motion, audio-only and human fallbacks.

Visible AI disclosure, SynthID verification if available, content policy, child-safety scope, export controls, monitoring and abuse response.

Cohort, traffic and duration ceilings; mandatory failures; fallback channel; owner; rollback trigger; model, price, language, endpoint or policy changes that force replay.

Primary sources

  1. Introducing Gemini 3.8 Live with Live AvatarGoogle · Reviewed: September 24, 2026 announcement date; Gemini Enterprise availability; multimodal conversation; asynchronous tool execution; 97-language claim; custom-avatar allowlist; SynthID; API link · Retrieved · Supports: Google announced Gemini 3.8 Live with Live Avatar on September 24, says it is available in Gemini Enterprise, and describes near-real-time audio and video, background tool calls, 97-language transitions, preset and allowlisted custom avatars, and SynthID watermarking. These capability and quality statements are vendor claims.
  2. Power your agents: Gemini 3.8 Live with Live Avatar is now generally availableGoogle Cloud · Reviewed: September 25, 2026 publication date; general availability; enterprise production positioning; features; US and EU endpoints; provisioned throughput; custom-avatar controls; SynthID; demos · Retrieved · Supports: Google Cloud calls Live Avatar generally available and ready for enterprise production in Gemini Enterprise. It lists web, mobile and kiosk use; asynchronous tools; live visual input; US and EU endpoints; provisioned throughput; curated avatars; allowlisted custom avatars; and SynthID. Gemini 3.8 Live Extended Thinking remains private preview.
  3. Gemini Live API overviewGoogle Cloud Documentation · Reviewed: Overview; key features; technical specifications; supported models; gemini-3.8-live availability and feature row; protocols and modalities; page update date · Retrieved · Supports: The documentation lists model ID gemini-3.8-live as generally available with native audio, transcriptions, voice activity detection, affective dialogue, proactive audio, tool use and Live Avatar. It documents a stateful WebSocket protocol, text/audio/image/video inputs, audio/text/avatar-video outputs, and 24 supported conversational languages. The page was last updated September 25, 2026 UTC.
  4. Generative AI on Agent Platform pricingGoogle Cloud · Reviewed: Gemini 3 standard-model pricing table; Gemini 3.8 Live API rows; modality-specific input and output prices; region; cache columns; billable-response rule · Retrieved · Supports: Google lists non-global Gemini 3.8 Live API prices per one million tokens: $0.75 text input, $1 image/video input, $3 audio input, $4.50 text output including response and reasoning, $12 audio output, and $1 avatar-video output. The table shows no cached-input rates and says only requests returning HTTP 200 are charged for input or output.
  5. Gemini 3.8 Audio model cardGoogle DeepMind · Reviewed: Publication date; model dependencies; inputs and outputs; context and output limits; distribution; intended use; known limitations; evaluation and Frontier Safety summaries · Retrieved · Supports: The model card says Gemini 3.8 Audio is based on Gemini 3 Pro. For Live it lists audio, image, video and text input with up to 128K context and 64K output; with Live Avatar it lists audio, video and text output with a 24K output limit. It names hallucination, occasional slowness or timeouts, a January 2025 knowledge cutoff and only a few minutes of continuous Live Avatar interaction as limitations.

Limitations

AccessAllGPT reviewed five first-party Google pages but had no Google Cloud credential, Gemini Enterprise tenant, provisioned throughput reservation or custom-avatar allowlist. We did not call gemini-3.8-live, open a WebSocket, generate audio or video, submit a likeness, execute a tool, inspect token accounting or billing, verify US/EU routing, test interruption or reconnection, measure latency or quality, reproduce the 97-language claim, evaluate any of the 24 documented languages, test the few-minutes duration boundary, inspect SynthID or reproduce Google’s safety evaluations. The reviewed sources are all vendor-authored. The launch pages’ 97-language statement and API overview’s 24-language statement remain unreconciled. Pricing, limits, endpoints, entitlements, documentation and behavior can change. This article is not a safety, privacy, identity-rights, accessibility, security or legal assessment.

Disclosures

AccessAllGPT did not receive early access, API credits, a Google Cloud account, an enterprise allowlist, product briefing, test data or compensation for this article. Google and Google DeepMind did not sponsor, review or endorse it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Google, Google DeepMind, OpenAI or organizations cited. Publication-wide relationships are listed on the disclosures page.

Further AccessAllGPT guidance

  1. Gemini 3.8 Flash: Same Rate, 40% Higher Cost in One Agent Suite
  2. Google Launches the Gemini App Globally on Windows 10 and 11
  3. Lyria 3.5 Is an $0.08 Full-Song API—but Its Java Samples Still Name Lyria 3
  4. Where Human Approval Belongs in AI Automation
  5. Design an Agent Benchmark That Predicts Production
  6. AI API Data Retention and Residency: Set the Procurement Gates
  7. AccessAllGPT Research methodology
  8. Publication disclosures

Continue the research

Get evidence-led updates for teams making production AI decisions.