Key takeaways
- AWS launched GLM 5.3 on Amazon Bedrock on October 5, 2026. Access is limited to eligible customers and requires either global.zai.glm-5.3 or us.zai.glm-5.3; in-Region on-demand inference is not supported.
- Global Standard list pricing is $1.68 input, $5.28 output, $0.312 cache read and $2.10 for a 30-minute cache write per million tokens. The US profile is 10% higher; Priority adds 75%, while Flex and Batch halve Standard rates.
- Bedrock documents text input and output, a 1M-token context window, 128K maximum output, and support through Responses, Chat Completions, InvokeModel and Converse.
- Implicit caching is on by default. Explicit caching requires at least 1,024 tokens per checkpoint and the model card lists retention of at least 30 minutes, but an eligible prefix does not guarantee a cache hit.
- Run a profile-specific shadow trial before migration. Pin API, profile, service tier and reasoning effort; measure accepted-task quality, total tokens, cache behavior, latency, tool outcomes and complete-task cost.
AWS launched GLM 5.3 on Bedrock on October 5
Amazon Web Services announced GLM 5.3 availability on Amazon Bedrock on October 5, 2026. This is a new managed route to Z.ai’s existing open-weight coding and agentic model, not a new GLM checkpoint. AWS says access is limited to eligible enterprise customers; teams without that entitlement must contact their AWS account team rather than assume the model will appear in every console.
The Bedrock model card marks the model active and documents text input and text output, a 1-million-token context window and up to 128,000 output tokens. The launch post positions the service for repository-scale coding, long-horizon agents and defensive security work. Those use cases are product positioning, not evidence that an unattended agent will remain correct or authorized over a million-token trajectory.
Requests must use a Global or US cross-Region profile
Applications must call global.zai.glm-5.3 or us.zai.glm-5.3. AWS lists zai.glm-5.3 as the base model ID, but explicitly says in-Region on-demand inference is unsupported. A request goes to a chosen source Region and Bedrock routes it for processing through the selected cross-Region inference profile.
That distinction belongs in the architecture record. Global CRIS gives AWS the broadest supported commercial-Region routing scope and carries the lower list price. The US profile keeps routing within its documented geography but is not the same as pinning processing to one Region. Verify the current destination set, account terms, logs and data requirements before sending regulated or location-restricted content.
Global Standard starts at $1.68 input and $5.28 output
The live Bedrock pricing table lists Global Standard rates per million tokens of $1.68 for input, $5.28 for output, $0.312 for cache reads and $2.10 for 30-minute cache writes. US CRIS costs $1.848 input, $5.808 output, $0.3432 cache read and $2.31 cache write—a 10% premium in each category.
AWS states that Priority pricing is 75% above Standard, while Flex and Batch are 50% below Standard. A tier multiplier is not a complete cost forecast: reasoning output, long agent trajectories, cache misses, repeated tool turns and retries can dominate a cheap input rate. Compare cost per accepted task and include service behavior, not only tokens from one successful call.
Four APIs do not mean four identical contracts
AWS supports the OpenAI-compatible Responses and Chat Completions APIs plus Bedrock InvokeModel and Converse. Its launch example uses the Responses API at the regional bedrock-runtime endpoint with global.zai.glm-5.3 and short-lived credentials generated from AWS CLI authentication. AWS recommends the OpenAI-compatible APIs for new applications because they expose the more complete feature set.
Treat each API surface as a separate integration until fixtures prove otherwise. Pin authentication, request fields, streaming event handling, structured output, function calls, reasoning controls, error mapping, retries and cancellation. A model-string replacement can still break parsers or authorization even when both providers describe an endpoint as OpenAI compatible.
Prompt caching has a price and an eligibility boundary
Implicit prompt caching is enabled by default. For explicit caching, AWS documents prompt_cache_options with cache breakpoints, a minimum of 1,024 tokens before each eligible checkpoint and retention of at least 30 minutes for this model. The pricing table separates cache reads from 30-minute cache writes, so both events must be visible in a cost model.
A breakpoint does not guarantee a hit. Stable prefixes, timing, routing and request construction still matter. During a trial, log input, cache-write, cache-read and output tokens separately; compare first-turn and repeated-turn latency; and test how prompt edits invalidate the prefix. Do not claim savings from the rate table without observed hit rates on the real conversation pattern.
The published model-size descriptions do not agree
AWS’s launch post calls GLM 5.3 a 753B-parameter mixture-of-experts model, matching the 753B label on Z.ai’s Hugging Face page. The live Bedrock model card instead says 744B total parameters and approximately 40B active per token. AccessAllGPT found no reconciliation on the reviewed pages, so procurement records should preserve both dated statements rather than silently choosing one.
The discrepancy does not change the Bedrock API price or profile ID, but it is a useful evidence-quality warning. Parameter count is not a performance result, and a hosted buyer does not operate those weights directly. Ask AWS or Z.ai to clarify the count if architecture scale is material to a disclosure, capacity comparison or technical assessment.
Coding and cyber scores remain vendor claims
Z.ai reports a 50% improvement over GLM 5.2 on its internal coding benchmark and publishes results including 28.3 on Terminal Bench 3.0, 66.9 on DeepSWE v1.1 and 84.5 on CyberGym. AWS repeats selected capability claims and demonstrates an authorized Strix security-testing workflow. AccessAllGPT did not reproduce any score or run that workflow; the numbers belong to the disclosed vendor harnesses and settings.
The model’s documented cyber capability raises the control requirement rather than granting authority. Any coding or security agent should run with scoped credentials, isolated execution, allowlisted targets, complete logs, step and spend ceilings, human approval before consequential action, and a kill switch. Only test systems you own or have explicit written permission to assess.
What builders and buyers should do next
Start by confirming account eligibility and selecting one cross-Region profile. Run a reversible shadow evaluation through the exact API and service tier intended for production. Freeze representative repository and tool-use tasks, set reasoning and output limits, and capture accepted outcomes, severe failures, latency, input, cache-write, cache-read and output tokens, tool calls, retries, refusals and cost.
Adopt only for task classes that clear predeclared quality, security, data-location and cost gates. Constrain the model to read-only analysis or patch proposals when execution authority is unnecessary. Wait when eligibility, routing or current pricing is not verified. Reject unattended consequential access without deterministic authorization, approval and rollback regardless of benchmark position.
Copy-ready GLM 5.3 Bedrock pilot record
Complete one record per profile, API, service tier and workload before production approval.
Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].
Trial, adopt, constrain, wait or reject; exact task, users, owner, consequence, review date and expiry.
Eligible account evidence, console availability, calling Region, IAM permissions and support contact.
global.zai.glm-5.3 or us.zai.glm-5.3; Responses, Chat Completions, InvokeModel or Converse; SDK, endpoint and authentication.
Reasoning effort, context and output caps, schemas, tools, streaming, stop conditions, retries and fallback.
Standard, Priority, Flex or Batch; dated input, output, cache-read and cache-write rates; complete-task budget.
Implicit or explicit mode, prefix and breakpoint design, token eligibility, observed hits, retention, invalidation and savings.
Cross-Region destination scope, data classes, residency evidence, retention, logs, contract and incident path.
Frozen tasks, baseline, repetitions, accepted outcomes, severe-error rule, latency, tokens, tools, retries and cost.
Scoped identity, sandbox, target authorization, approvals, spend and step limits, kill switch, rollback and re-test triggers.
Primary sources
Browse the publication-wide evidence index →
- Introducing GLM 5.3 on Amazon BedrockAmazon Web Services · Reviewed: October 5, 2026 publication date; eligibility statement; cross-Region profile IDs; supported APIs; authentication example; prompt-caching behavior; service tiers; Strix integration caveat; availability and cleanup notes · Retrieved · Supports: AWS announced GLM 5.3 on Amazon Bedrock for eligible enterprise customers, with Global and US cross-Region inference profiles, four API surfaces, implicit and explicit prompt caching, and Standard, Priority and Flex service tiers.
- GLM 5.3 — Amazon Bedrock model cardAmazon Bedrock Documentation · Reviewed: Model details; October 5 launch date; lifecycle; context and output limits; input and output modalities; APIs and endpoints; supported Bedrock features; caching minimum and retention; profile IDs; service tiers; regional availability; quotas and sample code · Retrieved · Supports: The live Bedrock model card documents a text-input/text-output model with a 1M-token context, 128K maximum output, profile IDs us.zai.glm-5.3 and global.zai.glm-5.3, cross-Region-only access, four API surfaces and at-least-30-minute explicit-cache retention.
- Amazon Bedrock pricing — Z AI GLM 5.3Amazon Web Services · Reviewed: GLM 5.3 On-Demand pricing table; Global and US CRIS input, output, cache-read and 30-minute cache-write rates; Standard tier basis; Priority premium; Flex and Batch discount · Retrieved · Supports: AWS lists Global Standard rates per million tokens of $1.68 input, $5.28 output, $0.312 cache read and $2.10 30-minute cache write; US CRIS rates are $1.848, $5.808, $0.3432 and $2.31 respectively. Priority adds 75%, while Flex and Batch are 50% below Standard.
- zai-org/GLM-5.3 model cardZ.ai on Hugging Face · Reviewed: Model overview; base-model relationship; benchmark table and run notes; 753B parameter label; coding and cyber claims; reasoning settings; deployment links; license label · Retrieved · Supports: Z.ai describes GLM-5.3 as a 753B-parameter open-weight model derived from the GLM-5.2 base through post-training, and reports coding and cyber benchmark results under disclosed but vendor-controlled harnesses. The card establishes publisher claims, not independently reproduced performance.
Limitations
AccessAllGPT reviewed public first-party pages but did not obtain an eligible AWS account; invoke zai.glm-5.3, global.zai.glm-5.3 or us.zai.glm-5.3; verify regional availability or destination routing; reproduce Z.ai or AWS benchmarks; measure quality, latency, throughput, reasoning, caching or cost; run the Strix example; inspect Bedrock safeguards; download or audit model weights; resolve the 744B-versus-753B documentation discrepancy; inspect training data; or perform a security, privacy, legal, residency or contractual audit. Prices, profile destinations, access criteria, features, quotas, model documentation and service tiers can change. A 1M context limit is capacity, not evidence of reliable recall or long-horizon autonomy.
Disclosures
AccessAllGPT received no AWS or Z.ai account, model access, credits, early information, briefing, demo, benchmark data, review or compensation for this article. Neither company sponsored, reviewed or endorsed it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Amazon Web Services, Z.ai, Hugging Face or Strix. Publication-wide relationships are listed on the disclosures page.
Further AccessAllGPT guidance
- GLM-5.3: Trial the Coding Gains, Contain the Cyber Capability
- Anthropic Says GLM-5.3 Built Working V8 Exploits in 50 of 410 Attempts
- GLM-5.3-Flash: Cheap, Open and Multimodal—Decide From the Artifact
- AWS Adds Grok 4.7 to Bedrock With Global and US Model IDs
- Move an LLM API Without Breaking Production
- AI API Data Retention and Residency: Set the Procurement Gates
- Design an Agent Benchmark That Predicts Production
- AccessAllGPT Research methodology
- Publication disclosures
Continue the research
Get evidence-led updates for teams making production AI decisions.