Key takeaways
- Anthropic announced Claude Sonnet 5.5 on September 28, 2026. The current Claude API model ID is claude-sonnet-5-5; Anthropic documents a 1-million-token context window, up to 128K output tokens, adaptive thinking, vision and tool use.
- Direct Claude API list pricing is $2 per million input tokens, $10 per million output tokens and $0.20 per million cache-read tokens. Anthropic says task-level cost can be up to 30% lower than Sonnet 5 because the new model may use fewer tokens; that is a vendor claim, not a guaranteed invoice reduction.
- AWS made the model available the same day on Amazon Bedrock through the Global CRIS inference profile under global.anthropic.claude-sonnet-5-5. AWS also says Claude Platform on AWS access is available in North America.
- This is not a drop-in name change for every Sonnet 5 integration. Anthropic documents five breaking changes involving between-tools thinking, forced tool use, thinking-block continuity, the accepted computer-use tool version and advisor-model compatibility.
- Run a shadow migration on complete tasks before switching production traffic. Compare accepted-task quality, latency, input, output and cache tokens, tool failures, response parsing and total cost rather than relying on vendor benchmark or cost-per-task claims.
Claude Sonnet 5.5 launched on September 28
Anthropic introduced Claude Sonnet 5.5 on September 28, 2026 as the second model in its Claude 5.5 family. The company positions it below Opus 5.5 for open-ended work requiring sustained judgment and above a future Haiku 5.5 for well-scoped coding, document, slide and spreadsheet tasks. The direct Claude API model ID is claude-sonnet-5-5.
The live model page documents a 1-million-token context window and a 128K-token maximum output. It supports text and image input, text output, multilingual use, vision, tool use and adaptive thinking. Those interface facts establish what developers can request; they do not establish reliable recall across the full context or safe autonomy across a long tool trajectory.
Direct API pricing is $2 in and $10 out per million tokens
Anthropic lists Claude Sonnet 5.5 at $2 per million input tokens, $10 per million output tokens and $0.20 per million cache-read tokens—the same rates it cites for Sonnet 5. The announcement says the new model costs up to 30% less per task in Anthropic’s testing because it typically uses fewer tokens for the same work. That task-economics result is a vendor claim and should not be substituted for a workload forecast.
A production comparison should price the entire request path: uncached and cached input, thinking and visible output, retries, failed tool turns and any surrounding retrieval or execution. A nominally unchanged token rate can still produce a different bill if the model changes output length, tool-call count, cache hit rate or retry behavior. AWS directs Bedrock customers to its live regional pricing table; confirm the rate displayed for the intended account and route before approval rather than assuming the direct Claude API tariff applies.
Amazon Bedrock has a separate global model ID
AWS announced Sonnet 5.5 availability on Amazon Bedrock on September 28. Its example uses global.anthropic.claude-sonnet-5-5 through the Global Cross-Region Inference, or Global CRIS, profile on bedrock-runtime. Developers can use Bedrock Invoke and Converse APIs or Anthropic’s Messages API against Bedrock. The example requires bedrock:InvokeModel and bedrock:InvokeModelWithResponseStream permissions.
Global CRIS is an access route, not a promise that every application can satisfy a fixed in-region processing requirement. AWS tells customers to consult the current supported-region documentation and separately says Claude Platform on AWS is available in North America. Confirm model access, inference-profile routing, data residency, quotas and contract terms in the actual account before moving regulated traffic.
Sonnet 5 migrations have documented breaking changes
Anthropic lists five breaking changes for applications already using Sonnet 5. Up-front thinking can be disabled with between_tools; forced tool use now returns an error; thinking blocks are tied to the model and conversation; an earlier computer-use tool version is not accepted on the Claude API and Google Cloud; and several older Claude models cannot serve as advisors for the advisor tool.
There is also a non-failing response-shape change: text between tool calls can arrive inside thinking blocks. An application that only streams ordinary text may appear silent between calls unless it chooses the documented display behavior or disables up-front between-tools thinking. Treat the migration as a parser, state and tool-contract change—not only a model string update.
Anthropic reports large coding gains and faster generation
Anthropic says Sonnet 5.5 generates output more than 30% faster than Sonnet 5. It reports 70.6% versus 10.3% on Terminal-Bench 4.0, 55.5% versus 34.1% on CursorBench 4.0 and 46.2% at Max effort versus 42.4% on FrontierCode 1.1 Main. It also reports a GDPval-AA v2.1 score within two points of Opus 5.5. These are vendor-selected and vendor-reported results under the announcement’s footnotes and settings.
The unusually large Terminal-Bench gap makes methodology review especially important. A buyer should not infer that its repository tasks improve by the same amount. Preserve the exact model, effort, harness, tools, time and token ceilings, scorer and task set when comparing results, then run representative private tasks that were not used to tune prompts or select the model.
Cyber safeguards now apply to a Sonnet model
Anthropic says Sonnet 5.5 is the first Sonnet model launched with cyber safeguards and fallbacks similar to those used for its most capable models because its cybersecurity capabilities are comparable to Opus 5. It says biology safeguards remain the same as Sonnet 5 and that both target a narrow set of high-risk requests. These are Anthropic’s descriptions of its own controls, not an independent safety certification.
Builders should test ordinary authorized security and life-sciences workflows for false refusals, escalation behavior and provider differences while maintaining their own authorization, monitoring and review controls. Provider safeguards do not grant permission to run a task, contain tool side effects or replace application-level access control.
Use a shadow migration before changing production traffic
Start by pinning the current Sonnet 5 baseline and Sonnet 5.5 candidate with the same prompts, tools, fixtures and effort policy. Replay complete workflows in shadow mode. Record accepted-task quality, first-token and completion latency, input, cache-read, thinking and output tokens, tool-call validity, parser failures, retries, refusals and cost per accepted task. Add explicit fixtures for every documented breaking change.
Move a small reversible cohort only after response parsing, tool state and cost gates pass. Keep the old model available during rollback. Wait when Global CRIS conflicts with residency requirements, an AWS account does not expose the expected route or price, or a tool integration depends on old thinking-block or computer-use behavior. Reject unattended authority where model output can trigger consequential writes without deterministic authorization and review.
Copy-ready Sonnet 5.5 migration record
Complete one record for each provider route, workload and authority level before replacing Sonnet 5.
Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].
Trial, deploy, constrain, wait or reject; exact task, user, owner, review date and expiry.
Claude API claude-sonnet-5-5 or Bedrock global.anthropic.claude-sonnet-5-5; account, region/profile, SDK and API version.
Context and output ceilings, effort, thinking policy, cache behavior, system prompt, tools, stop conditions and spend ceiling.
between_tools behavior, forced-tool errors, thinking-block continuity, computer-use tool version, advisor compatibility and streaming parser fixtures.
Frozen private tasks, baseline model, seeds, repetitions, scorers, acceptance threshold and documented benchmark relevance.
Quality, latency, uncached/cached input, thinking/output tokens, calls, retries, refusals and cost per attempted and accepted task.
Readable data, callable tools, write permissions, approvals, cyber or biology refusal tests, monitoring, kill switch and rollback.
Routing and residency, retention contract, logs, subprocessors, sensitive-data controls, deletion and incident process.
Account screenshots, price snapshot, request/response fixtures, test results, budget approval, rollout cohort, fallback and re-evaluation triggers.
Primary sources
Browse the publication-wide evidence index →
- Introducing Claude Sonnet 5.5Anthropic · Reviewed: September 28, 2026 announcement; positioning; performance, cost, speed and safety claims; benchmark table and footnotes; availability and getting-started guidance · Retrieved · Supports: Anthropic announced Claude Sonnet 5.5 on September 28, priced it at $2 per million input tokens, $10 per million output tokens and $0.20 per million cache-read tokens, and reported speed, cost-per-task, benchmark and safeguard claims that AccessAllGPT treats as vendor evidence.
- Claude Sonnet 5.5 model overviewClaude Platform Docs · Reviewed: Current Claude API model ID; context and output limits; token prices; supported platforms; adaptive thinking; effort control; breaking changes from Sonnet 5; tool and response-shape changes · Retrieved · Supports: Anthropic documents claude-sonnet-5-5 with a 1-million-token context window, 128K maximum output, $2 input and $10 output pricing per million tokens, and migration-sensitive changes to thinking blocks, forced tool use, computer use and advisor compatibility.
- Introducing Claude Sonnet 5.5 on AWSAmazon Web Services · Reviewed: September 28, 2026 availability announcement; Bedrock and Claude Platform on AWS access; workload positioning; model invocation example; IAM permissions; Global CRIS profile; regional caveats; monitoring and billing guidance · Retrieved · Supports: AWS says Sonnet 5.5 is available through Amazon Bedrock using global.anthropic.claude-sonnet-5-5 on the Global CRIS inference profile and through Claude Platform on AWS in North America. It documents Invoke, Converse and Anthropic Messages API routes.
- Anthropic models in Amazon BedrockAmazon Bedrock Documentation · Reviewed: Current Anthropic model catalog; Claude Sonnet 5.5 listing; model description; Bedrock-specific refusal billing notes; links to regional and feature support · Retrieved · Supports: The live Bedrock documentation lists Claude Sonnet 5.5 as an available Anthropic model and describes Bedrock-specific billing behavior for requests rejected before or during inference.
Limitations
AccessAllGPT reviewed four first-party Anthropic and AWS sources but did not have a Claude API or AWS credential for this work. We did not call Claude Sonnet 5.5; verify account-specific access, Global CRIS routing or regional pricing; reproduce Terminal-Bench, CursorBench, FrontierCode, GDPval-AA or other vendor results; compare Sonnet 5; measure speed, token use, cost per task, context reliability or 128K output behavior; test thinking, tool, computer-use or advisor migration changes; evaluate cyber or biology safeguards; inspect training data; or perform a security, privacy or residency audit. AWS’s public pricing component did not expose a stable account-and-region value during review, so this article does not assert that Bedrock pricing equals Anthropic’s direct API rate. IDs, prices, limits, routes, policies and availability can change.
Disclosures
AccessAllGPT did not receive Anthropic or AWS credentials, credits, early access, a briefing, benchmark data, review or compensation for this article. Anthropic and Amazon Web Services did not sponsor, review or endorse it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Anthropic, AWS or organizations cited. Publication-wide relationships are listed on the disclosures page.
Further AccessAllGPT guidance
- Claude Fable 5.1: Cache Economics Change the Migration Decision
- AWS Publishes an OpenAI-on-Bedrock Cost-per-Outcome Benchmark
- Move an LLM API Without Breaking Production
- Choose a Model Without Chasing the Leaderboard
- AI API Data Retention and Residency: Set the Procurement Gates
- Design an Agent Benchmark That Predicts Production
- AccessAllGPT Research methodology
- Publication disclosures
Continue the research
Get evidence-led updates for teams making production AI decisions.