Key takeaways

  • Google announced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, 2026. Both are available now to developers through the Gemini API and Google AI Studio under model IDs gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts.
  • Through December 31, 2026, both cost $0.50 per million text-input tokens; audio output is $9 per million tokens for Flash and $6 for Flash-Lite. Google says audio uses 25 tokens per second, equivalent to $0.00225 and $0.0015 per ten seconds. Every listed paid rate doubles on January 1, 2027.
  • Both models accept text only, return audio only, support single- and two-speaker scripts, and can use prebuilt, designed or replicated voices. The model card lists an 8K input limit and 64K output limit.
  • Voice replication requires reference audio plus a consent recording that matches the speaker. Google says generated audio receives SynthID; those controls do not replace an application’s own authority, disclosure, access, revocation and deletion rules.
  • Use Flash when expressive fidelity and broad language coverage matter most; test Flash-Lite first for high-volume work. Reprice any production forecast at the published January 2027 rates before approval.
01

Two Gemini TTS model IDs are available now

On September 23, 2026, Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The developer model IDs are gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts. Google’s launch page says both started rolling out that day in the Gemini API and Google AI Studio. Flash also appears in Gemini Notebook, while Flash-Lite appears in Google Vids. API access through Gemini Enterprise was described as coming soon rather than available now.

This is a current API launch, not only a research demo. The speech documentation provides request examples through the Interactions API and a supported-model matrix for both IDs. AccessAllGPT did not authenticate an account, however, so availability in a specific project, country, quota tier or enterprise contract remains to be confirmed in that account.

02

Flash and Flash-Lite share the API but target different jobs

Both models use the same schema, making the model ID the principal switch. Google positions Flash for maximum acoustic fidelity, acting control, difficult pronunciation, regional or minority dialects and long-form work. It positions Flash-Lite as the faster, lower-cost option for bulk production, read-aloud features and conversational voice-agent cascades. These workload claims come from Google and were not independently reproduced here.

The model card says both accept text up to 8K tokens and produce audio with an output ceiling of 64K tokens. The API documentation describes more than 130 supported languages for Flash and more than 100 for Flash-Lite. A language appearing in the table is not evidence that one voice, dialect, pronunciation set or safety behavior meets a production bar; test every required locale with the intended script distribution.

03

Launch pricing is temporary and doubles on January 1

Google’s pricing page lists the same text-input rate for both models through December 31, 2026: $0.50 per million tokens. Audio output costs $9 per million tokens for Flash and $6 for Flash-Lite. Google defines audio as 25 tokens per second and translates those rates to $0.00225 and $0.0015 per ten seconds respectively. The free tier is listed as free of charge, with free-tier data marked as used to improve Google products and paid-tier data marked as not used for that purpose.

The table already publishes the step-up: on January 1, 2027, text input becomes $1 per million tokens, Flash audio becomes $18 and Flash-Lite audio becomes $12. Input-cache and storage prices also double. A pilot approved on the launch tariff therefore needs two forecasts—one for the remaining 2026 window and one for steady-state 2027—plus retries, stored custom voices, preprocessing, delivery and application costs.

04

The output contract changes between unary and streaming requests

Unary requests return a complete 24 kHz mono WAV file with a RIFF header by default. Streaming requests instead return headerless 16-bit little-endian PCM chunks at 24 kHz. The API can also return linear PCM, μ-law or A-law and accepts an explicit sample rate. This matters during migration: code written for an older preview model may add a WAV header itself and corrupt an already wrapped Gemini 3.8 response.

For two-speaker audio, the documentation expects separate text items with speech_metadata identifying each speaker, and conversational mode for turn-taking. Sustained direction belongs in speech_metadata.style; momentary breaths, laughs or pauses belong as inline tags. Keep the transcript field verbatim. Add an automated format probe and listenable fixture before switching an existing production pipeline.

05

Voice design and replication are API resources, not just prompts

Gemini 3.8 TTS can use 30 named studio voices, an extended catalog, a designed voice created from a natural-language description, or a replicated voice created from recordings. Google’s documentation says stateful designed and replicated voices use voice_ IDs, share a 200-voice project quota and have a one-year retention period. Optional stateless replicated voicekey_ values are client-managed and expire after seven days.

Google says replication starts from a 30-second reference and requires a verbal consent recording from the same speaker. The launch post also says generated audio receives SynthID and replicated voices carry C2PA credentials. Those are vendor-described safeguards, not proof that every downstream player preserves provenance or that the requester owns the voice. Applications still need identity checks, scoped authority, visible synthetic-audio disclosure, restricted creation access, revocation and deletion.

06

Google reports strong voice rankings, but buyers need their own test

Google says Flash scored 71.4 on Hume AI’s Voice Design Benchmark and 60.8 for accent modeling, and says Flash and Flash-Lite ranked first and second on Hume’s Overall Quality Index. It also reports leading positions in selected languages on Voice Arena. These are vendor-selected and vendor-reported results. AccessAllGPT did not inspect evaluation samples, reproduce scores or establish that the tests represent a buyer’s voices and scripts.

A useful comparison fixes the transcript, voice authority, style instruction, output format and volume across Flash, Flash-Lite and the incumbent. Blind raters should score intelligibility, pronunciation, identity consistency, expressiveness and artifacts. Measure first-audio latency, completion latency, retries, billed tokens and accepted-audio minutes. Include difficult names, numbers, code switching, long passages, pauses and malformed markup—not just polished demos.

07

The practical decision is Flash, Flash-Lite or wait

Pilot Flash when a creative workflow needs designed characters, nuanced acting, long-form stability or the broader documented language set. Pilot Flash-Lite first when the main constraint is throughput or cost and the workload uses common languages with repeatable scripts. Because the schema is shared, route a frozen sample to both and choose on accepted output rather than the model name.

Wait when the required Gemini Enterprise API route, country, quota, voice-consent process or post-2026 budget is unresolved. Reject replication for any workflow that cannot prove speaker authority or revoke and delete a custom voice. Keep a non-cloned catalog voice or a text fallback ready. The API is available now; voice rights, workload quality and 2027 economics are still deployment decisions.

08

Copy-ready Gemini 3.8 TTS launch record

Complete one record for each model, voice-authority class, language set and delivery format before production traffic.

Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].

Pilot, deploy, constrain, wait or reject; user, channel, script type, monthly accepted minutes, owner, review date and expiry.

Project, region, gemini-3.8-flash-tts or gemini-3.8-flash-lite-tts, API/SDK version, quota, free or paid tier and enterprise availability.

Prebuilt, designed or replicated; speaker identity; reference and consent evidence; permitted uses; voice ID; storage mode; TTL; revocation and deletion.

Unary WAV or streaming PCM/telephony encoding, sample rate, speaker count, style metadata, allowed inline tags and client validation.

Frozen scripts and locales; blind-rater method; intelligibility, pronunciation, similarity, expressiveness, drift and artifact thresholds; results by model.

First-audio and completion latency, retries, failures, stream interruption, long-form continuity and load results at expected concurrency.

Text and audio tokens, launch-rate and January 2027 forecasts, cache/storage, retries and surrounding costs; cost per accepted audio minute and alert.

Synthetic disclosure, consent verification, SynthID/C2PA handling, creation access, impersonation blocks, monitoring, complaint process and incident owner.

Cohort, traffic ceiling, mandatory failures, non-cloned or text fallback, rollback owner and triggers for price, model, policy or format changes.

Primary sources

  1. Gemini 3.8 text-to-speech says helloGoogle · Reviewed: September 23, 2026 announcement; model positioning; voice design and replication; language and voice-library claims; performance controls; vendor-reported evaluations; safeguards; rollout by product · Retrieved · Supports: Google announced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, 2026. It says both began rolling out that day in the Gemini API and Google AI Studio, while Gemini Enterprise API access was still coming soon. Capability, quality and benchmark statements are Google claims.
  2. Text-to-speech generation (TTS)Google AI Studio Documentation · Reviewed: Model IDs; Interactions API request structure; single- and multi-speaker use; streaming; output formats; voice options and storage; supported languages; supported-model matrix; migration guide; limitations · Retrieved · Supports: The documentation identifies gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, documents text-only input and audio-only output, unary WAV and streaming PCM defaults, two-speaker generation, persistent and stateless custom voices, language coverage and migration changes from preview TTS models.
  3. Gemini API pricingGoogle AI Studio Documentation · Reviewed: Gemini 3.8 Flash TTS and Flash-Lite TTS standard pricing; free and paid tiers; temporary 2026 rates; January 2027 rates; context caching; audio-token conversion · Retrieved · Supports: Google lists paid prices through December 31, 2026 of $0.50 per million text-input tokens for both models, $9 per million audio-output tokens for Flash and $6 for Flash-Lite. It says prices double on January 1, 2027 and defines audio as 25 tokens per second.
  4. Gemini 3.8 Audio model cardGoogle DeepMind · Reviewed: September 15, 2026 publication date; model dependencies; inputs; outputs; distribution; evaluation approach; intended use; known limitations; safety summary · Retrieved · Supports: The model card says both TTS variants are based on Gemini 3 Pro, accept text up to 8K tokens and return audio with up to 64K output tokens. It documents distribution through the Gemini API and AI Studio and provides vendor-authored evaluation and limitation context.

Limitations

AccessAllGPT reviewed four first-party Google sources but had no Gemini API key, paid billing account, Gemini Enterprise tenant or voice-owner recording. We did not call either model, create or replicate a voice, verify consent matching, stream audio, inspect WAV or PCM output, test languages, compare voices, measure latency or quality, reproduce Hume or Voice Arena results, inspect SynthID or C2PA, confirm regional quota, test retention or deletion, or receive a bill. Pricing, quotas, languages, endpoints, model behavior and rollout status can change. The launch post says more than 100 languages broadly while the current documentation gives different counts for Flash and Flash-Lite; this article uses the documentation for model-specific planning. This is not a biometric, copyright, privacy, accessibility, security or legal assessment.

Disclosures

AccessAllGPT did not receive Google credits, an API account, early access, custom voices, a product briefing, test data, review or compensation for this article. Google and Google DeepMind did not sponsor, review or endorse it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Google, Google DeepMind, OpenAI or organizations cited. Publication-wide relationships are listed on the disclosures page.

Further AccessAllGPT guidance

  1. Gemini 3.8 Live Avatar Is Generally Available—Here Are the API Price and Limits
  2. AWS Adds Qwen3-TTS Voice Cloning to SageMaker JumpStart
  3. Gemini 3.8 Flash: Same Rate, 40% Higher Cost in One Agent Suite
  4. AI API Data Retention and Residency: Set the Procurement Gates
  5. Choose a Model Without Chasing the Leaderboard
  6. AccessAllGPT Research methodology
  7. Publication disclosures

Continue the research

Get evidence-led updates for teams making production AI decisions.