Key takeaways

  • OpenAI announced GPT-6.1 Sol on September 29, 2026. The current API model ID is gpt-6.1-sol, and OpenAI positions it for complex coding, computer use and professional work between GPT-6 Astra and the lower-cost GPT-6 Luna.
  • Standard short-context list pricing is $2 per million input tokens, $0.10 cached input, $2.50 cache writes and $10 output. That is one-fifth of Astra’s $10 input and $50 output rates, but it is not a one-fifth cost-per-task guarantee.
  • The model accepts text and images, returns text, and documents a 1,050,000-token context window, 922,000 maximum input and 128,000 maximum output. Requests above 272K input tokens apply higher rates to the full request.
  • Tool calling requires the Responses API. Chat Completions is supported without tools; reasoning effort can be low, medium, high, xhigh or max, with medium as the default.
  • Run a workload-specific comparison against Astra and the current production model. Pin effort, API surface, context band and processing mode, then compare accepted-task quality, latency, token use, tool outcomes and total cost.
01

GPT-6.1 Sol launched on September 29

OpenAI published GPT-6.1 Sol on September 29, 2026 at 10:00 UTC. Its official feed describes the release as near-Astra intelligence for coding, computer use and professional work at one-fifth of Astra’s standard API input and output token prices. The live API identifier is gpt-6.1-sol.

OpenAI’s current model guide places Sol between GPT-6 Astra, its highest-intelligence option, and GPT-6 Luna, its fastest and lowest-cost family member. “Near-Astra performance” is vendor positioning, not an independent result. The accessible first-party documentation does not establish how close the models are on a reader’s workload, and AccessAllGPT did not run either model.

02

The headline rate is $2 in and $10 out

For short-context Standard requests, OpenAI lists $2 per million input tokens, $0.10 per million cached-input tokens, $2.50 per million cache-write tokens and $10 per million output tokens. GPT-6 Astra is listed at $10 input and $50 output, which supports OpenAI’s one-fifth token-rate comparison. It does not prove one-fifth the cost for a completed task because models may use different reasoning, output, retry and tool-call budgets.

Fast mode costs twice Standard. Batch and Flex are 50% below Standard, and regional processing adds 10% where available. Cache writes cost 1.25 times uncached input while cache reads cost 5% of uncached input. Record the actual processing mode, region and cache behavior in any cost comparison rather than quoting the base rate alone.

03

Long prompts move the entire request into a higher price band

GPT-6.1 Sol has a 1,050,000-token context window, with up to 922,000 input tokens and 128,000 output tokens. For a request containing more than 272,000 input tokens, OpenAI says input and cache rates double and output rises by 50% for the full request. The resulting Standard rates are $4 input, $0.20 cached input, $5 cache writes and $15 output per million tokens.

That threshold creates a discontinuity, not a marginal surcharge only on tokens above 272K. Retrieval and agent systems should measure the complete assembled request after instructions, history, documents and tool results are included. A large context ceiling is capacity, not evidence of reliable fact use across every position or of economical full-corpus prompting.

04

Tool use belongs on the Responses API

OpenAI documents Responses and Chat Completions support, but tool calling is available only through Responses. The supported Responses tools include web search, file search, image generation, Code Interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search. Chat Completions remains available for requests without tools.

Reasoning effort supports low, medium, high, xhigh and max; medium is the default, while none and minimal are unsupported. An evaluation that omits effort is therefore evaluating medium—not a neutral no-reasoning baseline. Pin effort explicitly and cap output, tool calls, elapsed time and spend when comparing quality or cost.

05

Residency and endpoint limits narrow the deployment choices

The model reference says GPT-6.1 Sol supports US and EU data residency, but Fast mode is unavailable with EU residency. Regional processing carries the documented 10% premium where offered. Eligibility and contractual controls remain account-specific; the presence of an EU option on a model page is not by itself evidence that a particular organization’s retention, subprocessors or residency requirements are satisfied.

The model accepts text and image input and produces text. It does not support audio or video input, fine-tuning, Realtime, Assistants or predicted outputs. Batch is supported. These interface boundaries matter during migration: a system using tool-enabled Chat Completions, a fine-tuned model or Fast mode with EU residency cannot switch by changing only the model string.

06

Compare complete tasks before replacing Astra or GPT-6 Sol

Build a frozen set of representative coding, computer-use and professional tasks. Run the current model, GPT-6.1 Sol and Astra with explicit effort, the same tools and equivalent stop conditions. Score accepted outcomes, not benchmark impressions. Capture first-token and completion latency, input by price band, cache writes and reads, reasoning and visible output, tool calls, retries, refusals and cost per accepted task.

Pilot when the lower token rate could materially change economics and the Responses API fits the tool contract. Wait when account access, residency eligibility, long-context behavior or migration requirements are unverified. Reject unattended consequential actions without deterministic authorization, least privilege, approval and rollback. The immediate decision is whether GPT-6.1 Sol earns a bounded comparison—not whether “near-Astra” settles the production choice.

07

Copy-ready GPT-6.1 Sol evaluation record

Complete one record for each workload, context band, processing mode and residency route before changing production traffic.

Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].

Trial, deploy, constrain, wait or reject; exact task, user, owner, review date and expiry.

gpt-6.1-sol; Responses or tool-free Chat Completions; snapshot, modalities, schema, tools and fallback.

Explicit low, medium, high, xhigh or max effort; input, output, tool-call, elapsed-time and spend ceilings.

Input at or below 272K or above 272K; Standard, Fast, Batch or Flex; region premium and dated token rates.

Write and read tokens, reusable-prefix policy, invalidation, hit rate and total cost per attempted and accepted task.

Frozen private tasks, Astra and production baselines, repetitions, scorers, acceptance gates and missing-run handling.

Quality, latency, token classes, tool outcomes, retries, malformed responses, refusals, rate limits and failure recovery.

US or EU route, eligibility evidence, retention, sensitive data, permissions, approvals, kill switch and audit trail.

Account access, price snapshot, request and response fixtures, test results, rollout cohort, fallback and re-evaluation triggers.

Primary sources

  1. Introducing GPT-6.1 SolOpenAI · Reviewed: OpenAI News RSS entry; title; description; canonical URL; Product category; September 29, 2026 publication timestamp · Retrieved · Supports: OpenAI’s official feed records the GPT-6.1 Sol launch at 10:00 UTC on September 29, 2026 and describes the model as intended for coding, computer use and professional work at one-fifth of GPT-6 Astra’s standard API input and output token prices.
  2. GPT-6.1 Sol model referenceOpenAI Developers · Reviewed: Model ID and positioning; modalities; context and output limits; knowledge cutoff; reasoning-effort levels; endpoints; tools; data residency; token prices; long-context, processing-mode and regional premiums; rate limits · Retrieved · Supports: The live model reference documents gpt-6.1-sol with text and image input, text output, a 1,050,000-token context window, 922,000 maximum input, 128,000 maximum output, five supported reasoning efforts, Responses API tool calling, US and EU data residency, and current Standard token prices.
  3. OpenAI API pricingOpenAI Developers · Reviewed: Standard short- and long-context token tables for GPT-6 Astra, GPT-6.1 Sol and GPT-6 Sol; Batch and Flex pricing; cache-write pricing; processing and regional-price qualifications · Retrieved · Supports: OpenAI’s pricing table lists GPT-6.1 Sol at $2 input, $0.10 cached input, $2.50 cache writes and $10 output per million short-context tokens, rising to $4, $0.20, $5 and $15 respectively for requests above the documented long-context threshold.
  4. Latest models guideOpenAI Developers · Reviewed: GPT-6 family selection; GPT-6.1 Sol positioning; Responses API guidance; reasoning-effort options; GPT-6 Sol migration notice; Astra comparison language · Retrieved · Supports: OpenAI positions GPT-6.1 Sol between GPT-6 Astra and GPT-6 Luna, recommends the Responses API for tool use, and tells existing GPT-6 Sol users to review migration guidance before switching rather than treating the model as a drop-in replacement.

Limitations

The canonical OpenAI launch page was protected by a Cloudflare JavaScript-and-cookie challenge during review. AccessAllGPT verified the title, canonical URL, Product category, description and September 29 publication timestamp through OpenAI’s official News RSS feed, then checked the live developer model page, pricing tables and model guide directly. We did not call gpt-6.1-sol; obtain an account-specific price or residency determination; reproduce any vendor evaluation; compare GPT-6 Astra, GPT-6 Sol or GPT-6 Luna; measure quality, latency, context recall, reasoning, computer use, tool behavior, caching, rate limits or cost; test safety controls; inspect training data; or conduct a security, privacy or contractual audit. The launch page access gap means this article intentionally omits unverified benchmark and safety claims. IDs, prices, thresholds, limits, tools, policies and availability can change.

Disclosures

AccessAllGPT did not receive OpenAI credentials, credits, early access, a briefing, benchmark data, review or compensation for this article. OpenAI did not sponsor, review or endorse it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with OpenAI. Publication-wide relationships are listed on the disclosures page.

Further AccessAllGPT guidance

  1. GPT-6 Astra: Buy Autonomy Only Where Control Improves
  2. GPT-5.6 Sol Ultrafast: Pay for the Tail-Latency SLA, Not the Label
  3. AWS Publishes an OpenAI-on-Bedrock Cost-per-Outcome Benchmark
  4. Move an LLM API Without Breaking Production
  5. Choose a Model Without Chasing the Leaderboard
  6. AI API Data Retention and Residency: Set the Procurement Gates
  7. Design an Agent Benchmark That Predicts Production
  8. AccessAllGPT Research methodology
  9. Publication disclosures

Continue the research

Get evidence-led updates for teams making production AI decisions.