Key takeaways
- AWS announced on September 25, 2026 that Qwen3-TTS-12Hz-1.7B-Base is deployable from SageMaker JumpStart as a managed real-time endpoint. AWS also lists the CustomVoice TTS and Qwen3-ASR 1.7B options in JumpStart.
- The worked deployment pins model ID huggingface-ttsvoiceclone-qwen3-tts-12hz-1-7b-base, version 1.0.1, on ml.g6.4xlarge: one NVIDIA L4 with 24 GB of GPU memory. It sets SM_VLLM_GPU_MEMORY_UTILIZATION to 0.45.
- Voice cloning requires a reference clip and transcript. The endpoint expects an OpenAI-style speech payload, task_type Base, and CustomAttributes set to route=/v1/audio/speech; the example generates 24 kHz mono WAV.
- Qwen lists ten languages and Apache-2.0 weights. Qwen’s “three-second rapid voice clone” and low-latency statements are vendor claims, not results reproduced by AccessAllGPT.
- There is no bundled per-character or per-token price in the announcement. Budget the regional GPU endpoint, instance count, uptime, storage, monitoring and surrounding application—and obtain explicit authority to clone every voice.
AWS published the JumpStart path on September 25
On September 25, 2026, AWS published a production-oriented path for deploying Qwen3-TTS-12Hz-1.7B-Base from SageMaker JumpStart to a managed real-time inference endpoint. This is an availability and serving update, not a new Qwen checkpoint: the model family and downloadable weights already existed. What changed is a pre-built JumpStart model artifact and serving container that AWS customers can deploy without writing a custom inference handler.
The exact JumpStart model ID in the walkthrough is huggingface-ttsvoiceclone-qwen3-tts-12hz-1-7b-base, pinned at model version 1.0.1. AWS says Qwen3-TTS-12Hz-1.7B-CustomVoice and Qwen3-ASR-1.7B are also available through JumpStart. AccessAllGPT did not authenticate to AWS, search a regional JumpStart catalog or verify quota, so teams should confirm that the model and supported instance are visible in their intended account and region before planning a migration.
The reference deployment uses one 24 GB L4 GPU
AWS deploys the 1.7B Base model on ml.g6.4xlarge, described in the post as one NVIDIA L4 GPU with 24 GB of memory. The container runs a talker stage and a code2wav stage on that GPU. The required example setting SM_VLLM_GPU_MEMORY_UTILIZATION=0.45 gives each stage the same reservation; AWS explains that 0.45 plus 0.45 leaves approximately ten percent of the GPU outside those reservations.
From its startup logs, AWS reports 3.66 GiB for talker weights, 0.45 GiB for code2wav weights, 6.08 GiB of reserved talker KV cache and a 56,928-token KV-cache budget. Those figures demonstrate that AWS’s example started with headroom; they do not establish concurrency, tail latency or capacity for another prompt distribution. Measure request size, audio duration, queueing and failure rate on the exact model version and instance before setting autoscaling targets.
The endpoint contract has two easy-to-miss requirements
The request uses an OpenAI-style speech payload with task_type set to Base. It includes target text, a base64 data URI for the reference audio, the reference transcript, language and response format. AWS recommends converting the reference to 24 kHz mono WAV; the demonstrated response is also 24 kHz mono WAV.
SageMaker exposes one /invocations path, but the container routes that path to a completions handler by default. AWS says callers must send CustomAttributes="route=/v1/audio/speech" to reach text-to-speech; omitting it causes rejection. Treat the model ID, version, route header, schema and audio normalization as one versioned client contract. Add a startup canary that synthesizes authorized test audio and fails deployment if content type, sample rate or route behavior changes.
Qwen documents ten languages and downloadable Apache-2.0 weights
The Qwen model card lists Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish and Italian. It labels Qwen3-TTS-12Hz-1.7B-Base as streaming-capable, suitable as a fine-tuning base and capable of cloning from three seconds of user audio. The Hugging Face repository labels the model Apache-2.0 and exposes Safetensors weights.
Those facts give buyers an exit path that a closed speech API may not: the underlying checkpoint can be downloaded and served outside JumpStart subject to the license and dependencies. Portability is not equivalence, however. AWS’s pre-built container, model version, vLLM-Omni integration and custom route are a distinct serving product. The model card says vLLM-Omni online serving is still planned, while AWS’s container exposes a managed online endpoint; do not assume every local recipe reproduces that contract.
Cross-lingual cloning is available, but quality is still a claim
AWS demonstrates reusing an English reference clip to generate Chinese speech while preserving the speaker identity. It recommends one request per sentence or paragraph for longer content so the voice remains consistent. Qwen describes low-latency streaming and a 97 ms minimum end-to-end synthesis result in its model card. These are vendor demonstrations and claims, not AccessAllGPT measurements.
A useful acceptance set pairs every required reference language with every output language and includes names, numbers, abbreviations, code switching, noisy recordings, accents and long-form continuity. Blind raters should score intelligibility, speaker similarity, pronunciation and naturalness, while a separate review checks whether the output could mislead listeners. Do not infer ten-language production quality from a supported-language list.
Pricing is infrastructure-based, not bundled per character
The AWS launch post does not quote a Qwen3-TTS per-character, per-second or per-token rate. SageMaker AI pricing says customers pay for what they use through on-demand pricing or Savings Plans. For this deployment, the practical base unit is the regional ml.g6.4xlarge endpoint and its running time, then additional instances under scaling. The generic SageMaker free-tier row covers m4.xlarge or m5.xlarge Real-Time Inference hours, not the L4 GPU instance in this walkthrough.
Price one steady endpoint, expected scaling hours, storage, logs, data transfer, retries and preprocessing in the intended region. Divide the full monthly estimate by accepted audio minutes—not raw requests—and compare that with a managed speech API under the same quality and availability gates. Delete the endpoint, endpoint configuration and model after evaluation; AWS explicitly includes cleanup commands because an idle real-time endpoint continues to create cost.
Voice authority is the production gate
A few-second cloning capability makes consent and impersonation controls part of the minimum architecture. Require documented authority for each reference voice and each use, preserve the source and permission record, restrict who can submit references, and bind reusable voice representations to an owner, purpose and expiration. Block public-figure, employee, customer and child voices unless the organization has an explicit lawful and ethical basis reviewed by the appropriate owner.
Disclose synthetic speech to listeners where context could create confusion. Add rate limits, abuse monitoring, deletion and revocation, and a response path for unauthorized clones. Keep reference audio and transcripts out of application logs, and verify encryption, retention, access and deletion in the actual AWS configuration. “Data stays within your AWS account” is AWS’s architectural claim; account ownership alone does not prove a private subnet, correct IAM policy or compliant retention setup.
The decision: pilot the endpoint with one authorized voice
Pilot this JumpStart path when the team needs controllable Qwen weights inside its AWS environment, has L4 quota, can hold one narrow language pair and voice to measurable acceptance criteria, and prefers infrastructure pricing over a bundled speech tariff. Pin version 1.0.1, preserve the exact route and schema, start with one instance, and load-test to a cost per accepted audio minute before enabling autoscaling.
Wait when the region, quota, instance price, container support or voice authority is unresolved. Reject deployment when references can be supplied without consent, generated audio can impersonate a person without disclosure, or the application cannot revoke a voice and delete its artifacts. JumpStart reduces serving work; it does not turn voice cloning into an ordinary stateless text endpoint.
Copy-ready Qwen3-TTS endpoint launch record
Complete one record for each model version, region, voice-authority class and language set before production traffic.
Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].
Pilot, deploy, constrain, wait or reject; use case; accountable product, ML, security and voice-rights owners; review and expiry dates.
AWS account, region, JumpStart model ID, model version 1.0.1, container digest if exposed, ml.g6.4xlarge quota, instance count and scaling policy.
Reference-speaker identity, consent evidence, permitted content and channels, expiration, revocation, deletion and prohibited identities.
24 kHz mono reference, transcript, Base task type, route=/v1/audio/speech, allowed languages, response format, size limits and rejection behavior.
Frozen texts and references; speaker similarity, intelligibility, pronunciation, long-form continuity, cross-lingual and adverse-input results; rater method.
Startup, first-audio and completion latency; concurrency; 4XX/5XX; GPU and memory utilization; queueing; autoscaling and overload behavior.
Regional instance rate, base and scaled hours, storage, logs, transfer, retries and preprocessing; accepted audio minutes; cost per accepted minute and budget alert.
Reference and transcript storage, IAM, network, encryption, logging exclusions, synthetic disclosure, rate limit, monitoring, complaint and incident process.
Cohort, traffic ceiling, canary, mandatory failures, endpoint deletion test, fallback provider or mode, rollback owner and replay triggers.
Primary sources
Browse the publication-wide evidence index →
- Deploying real-time personalized speech with Qwen3-TTS on Amazon SageMaker AIAmazon Web Services · Reviewed: September 25, 2026 publication date; JumpStart availability; model and container identifiers; architecture; prerequisites; deployment configuration; GPU-memory observations; request contract; languages; output format; scaling; CloudWatch monitoring; cleanup · Retrieved · Supports: AWS says Qwen3-TTS-12Hz-1.7B-Base can now be deployed from SageMaker JumpStart to a managed real-time endpoint. Its worked configuration pins JumpStart model ID huggingface-ttsvoiceclone-qwen3-tts-12hz-1-7b-base at version 1.0.1 on ml.g6.4xlarge, sets GPU memory utilization to 0.45, requires a custom TTS route attribute, and returns 24 kHz audio. AWS reports startup memory and KV-cache observations from its example; these are vendor measurements, not AccessAllGPT benchmarks.
- Qwen3-TTS-12Hz-1.7B-Base model cardQwen on Hugging Face · Reviewed: License; model family overview; released-model table; language support; streaming status; three-second clone description; local installation and inference; reference-audio contract; vLLM status; evaluation methodology; file metadata · Retrieved · Supports: The Qwen model card labels the weights Apache-2.0, lists ten supported languages and describes the 1.7B Base checkpoint as capable of a three-second rapid voice clone and fine-tuning. It documents reference audio plus transcript as the normal cloning inputs. The card’s latency, quality and benchmark statements are Qwen claims; it also says vLLM-Omni online serving is planned later even though AWS provides its own managed serving container.
- Amazon SageMaker AI pricingAmazon Web Services · Reviewed: Pricing overview; on-demand and Savings Plans choices; pay-for-use statement; Real-Time Inference free-tier row; pricing calculator link · Retrieved · Supports: AWS describes SageMaker AI as usage-priced, with on-demand pricing that has no minimum fee or upfront commitment and optional Savings Plans. The page does not publish one universal Qwen3-TTS per-character or per-token tariff; endpoint cost depends on infrastructure and region. Its Real-Time Inference free-tier row names older m4.xlarge or m5.xlarge instances, not the ml.g6.4xlarge GPU used in the Qwen walkthrough.
Limitations
AccessAllGPT did not have an AWS account, SageMaker GPU quota, JumpStart catalog view or billing access. We did not deploy model version 1.0.1, accept its EULA, inspect the container or weights, submit reference audio, clone a voice, validate consent controls, call the custom route, measure startup, latency, throughput, memory, quality or cross-lingual behavior, trigger autoscaling, inspect CloudWatch, calculate a region-specific ml.g6.4xlarge price or test cleanup. AWS’s memory values are observations from its walkthrough; Qwen’s three-second cloning, quality, language and latency statements are vendor claims. The reviewed sources do not provide a bundled model tariff, universal concurrency figure, regional availability matrix for this exact JumpStart entry, watermark guarantee or complete voice-misuse policy. Models, containers, prices, regions, quotas and documentation can change. This article is not a license, privacy, biometric, consent, accessibility, security or legal assessment.
Disclosures
AccessAllGPT did not receive AWS credits, an account, Qwen access, model files, a product briefing, test data, review or compensation for this article. AWS, Alibaba Cloud, Qwen and Hugging Face did not sponsor, review or endorse it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Amazon Web Services, Alibaba Cloud, Qwen, Hugging Face, OpenAI or organizations cited. Publication-wide relationships are listed on the disclosures page.
Further AccessAllGPT guidance
- Gemini 3.8 Live Avatar Is Generally Available—Here Are the API Price and Limits
- GLM-5.3 Flash Ships Open Weights
- WeatherNext 3 Is Operational—but Managed Inference Still Specifies WeatherNext 2
- AI API Data Retention and Residency: Set the Procurement Gates
- Choose a Model Without Chasing the Leaderboard
- Design an Agent Benchmark That Predicts Production
- AccessAllGPT Research methodology
- Publication disclosures
Continue the research
Get evidence-led updates for teams making production AI decisions.