Key takeaways

  • OpenAI published “A model guide for the GPT-6 family” on October 2, 2026 at 16:15 UTC. Its official feed says the guide covers model choice, reasoning effort, prompts and skills, tool coordination and production preparation.
  • The current API catalog names gpt-6-astra for complex reasoning and coding, gpt-6.1-sol for balancing intelligence and cost, and gpt-6-luna for cost-sensitive high-volume work. These are vendor positions, not independent rankings.
  • Luna’s short-context Standard list price is $0.10 input, $0.01 cached input, $0.125 cache writes and $0.50 output per million tokens. Astra’s corresponding $10 input and $50 output rates create a 100× token-rate spread.
  • All three featured models document a 1,050,000-token context window and 128,000-token maximum output, but a request above 272K input tokens moves its full token bill to the higher long-context band.
  • Do not route by label alone. Compare the same frozen tasks across explicit model, reasoning effort, processing mode and tool contract, then choose the least expensive configuration that clears a predeclared quality and safety gate.
01

OpenAI published the GPT-6 family guide on October 2

OpenAI’s official News feed records “A model guide for the GPT-6 family” at 16:15 UTC on October 2, 2026. The feed describes guidance for startups on choosing GPT-6 models, tuning reasoning effort, improving prompts and skills, coordinating tools and preparing workflows for production. The canonical page was behind a Cloudflare JavaScript-and-cookie challenge during review, so AccessAllGPT is not attributing unverified examples or benchmark claims to it.

The directly accessible developer catalog supplies the current selection headline: start with GPT-6 Astra for complex reasoning and coding, choose GPT-6.1 Sol to balance intelligence and cost, or use GPT-6 Luna for cost-sensitive high-volume workloads. All three accept text and image input and return text through the Responses API. This is a product-family map, not evidence that one choice wins a particular workload.

02

Luna is the low-cost endpoint in the current family

The live Luna reference identifies gpt-6-luna and lists Standard short-context prices of $0.10 per million input tokens, $0.01 cached input, $0.125 cache writes and $0.50 output. GPT-6.1 Sol is listed at $2 input and $10 output, while Astra is $10 input and $50 output. On those two headline categories, Luna is one-twentieth of Sol and one-hundredth of Astra.

Token-rate ratios are not cost-per-task results. A cheaper model may require more reasoning tokens, retries, validation or escalation, while a more expensive model may be unnecessary for bounded extraction. Tool calls can add separate fees. Budgeting should therefore use accepted outcomes from the intended workflow, not a multiplication of one advertised rate by an assumed constant token count.

03

The shared 1.05M context does not make the models interchangeable

Astra, GPT-6.1 Sol and Luna each document a 1,050,000-token context window and 128,000-token maximum output; Luna and Sol specify up to 922,000 input tokens. Luna accepts text and images and returns text. Its May 18, 2026 knowledge cutoff is newer than the April 30 date shown for Astra and GPT-6.1 Sol, but a cutoff date does not measure reasoning quality, factual reliability or coverage of any specific source.

For all three, requests above 272,000 input tokens move the full request to long-context pricing. Luna then costs $0.20 input, $0.02 cached input, $0.25 cache writes and $0.75 output per million tokens. Large context is capacity, not a reason to send an entire corpus without testing retrieval quality, evidence position, latency and total cost.

04

Reasoning effort is part of the model choice

Luna supports none, low, medium, high, xhigh and max reasoning effort, with medium as the default. GPT-6.1 Sol and Astra support low through max but not none. OpenAI’s selection page places Luna Low on fine-grained edits, well-scoped problems and simple extraction; Luna Extra high on constrained multi-application work; Sol on complex and revisable deliverables; and Astra on demanding work. OpenAI explicitly calls this ladder a starting point and recommends experimentation.

Every comparison should set effort explicitly. Otherwise a Luna request defaults to medium, and two teams can report different results while believing they tested the same model. Record model ID, effort, prompt version, tools, input and output ceilings, processing mode and stopping policy together as one configuration.

05

Responses is the complete tool surface

For Luna, OpenAI lists Responses and Chat Completions as supported endpoints, but says to use Responses for built-in tools and function calling. Chat Completions function calling is limited to reasoning_effort none. Through Responses, the documented tool list includes web and file search, image generation, Code Interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search.

A migration that changes only the model string can therefore alter or break the tool contract. Luna does not support Realtime, Live, Assistants or fine-tuning. Before routing production traffic, replay schemas, forced-tool cases, parallel calls, refusals, timeouts, retries and approval paths on the exact endpoint and effort level intended for deployment.

06

Processing mode and residency change the base price

OpenAI lists Batch and Flex at half of Standard rates and Fast at twice the applicable rate. For short-context Luna, that makes Batch or Flex $0.05 input and $0.25 output, versus Fast at $0.20 and $1.00 per million tokens. These tiers trade scheduling or service behavior as well as price; do not substitute one without checking its latency and operational contract.

The Luna reference says EU data residency is available with Standard, Flex and Batch processing, subject to eligibility. It does not list Fast for EU residency. Regional processing adds a 10% premium where available. A public model-page statement is not account-specific evidence for retention, subprocessor, residency or contract requirements; capture the project setting and applicable terms before sending regulated data.

07

Default rate limits support scale tests, not scale guarantees

OpenAI’s Luna reference lists default limits from 500 requests per minute and 500,000 tokens per minute at usage tier 1 to 30,000 RPM and 180 million TPM at tier 5, with separate Batch queue limits. These are tiered service limits, not promised application throughput or evidence that an account currently has tier-5 capacity.

A high-volume pilot should model bursts, cache misses, long-context crossings, tool latency, retries and provider errors. Measure queueing and throttling under a production-like arrival pattern. Add a bounded fallback policy rather than silently escalating every failure to Sol or Astra, where the same token volume can change spend by orders of magnitude.

08

Choose the lightest configuration that clears the gate

Freeze representative tasks and compare Luna, GPT-6.1 Sol and Astra with explicit effort, the same usable evidence and equivalent tool permissions. Score accepted outcomes, severe factual or action errors, latency percentiles, input by context band, cache writes and reads, reasoning and visible output, tool calls, retries, reviewer time and total cost per accepted task. Predeclare non-compensable safety failures before looking at aggregate scores.

Adopt Luna for bounded, high-volume work only when it clears the workload gate and its endpoint, residency and rate limits fit. Escalate selected cases to Sol or Astra when measured quality—not the label—justifies the price. Wait when access or data controls are unverified. Reject any unattended consequential workflow that lacks deterministic authorization, least privilege, approval and rollback regardless of which GPT-6 family member generates the proposal.

09

Copy-ready GPT-6 family routing record

Complete one record per workload before routing between Luna, GPT-6.1 Sol and Astra.

Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].

Exact task, users, owner, frequency, consequence; adopt, constrain, wait or reject; review date and expiry.

gpt-6-luna, gpt-6.1-sol or gpt-6-astra; snapshot, explicit reasoning effort, endpoint, schema, tools and stop conditions.

Requests and tokens by hour, concurrency, input distribution, 272K crossings, output caps and Batch queue needs.

Standard, Batch, Flex or Fast; short or long context; region premium; tool fees; dated rates and account terms.

Frozen cases, accepted-outcome rubric, severe-error rule, repetitions, evaluator agreement and missing-run handling.

Latency percentiles, rate limits, throttles, tool outcomes, retries, refusals, fallbacks and cost per accepted task.

Residency eligibility, retention, sensitive inputs, permissions, approval boundary, audit trail, kill switch and rollback.

Default configuration, measurable escalation criteria, maximum spend, no-silent-upgrade rule and fallback behavior.

Dated docs, project settings, test traces, score outputs, token logs, invoice sample, approver and re-evaluation triggers.

Primary sources

  1. A model guide for the GPT-6 familyOpenAI · Reviewed: OpenAI News RSS entry; title; description; canonical URL; Product category; October 2, 2026 publication timestamp · Retrieved · Supports: OpenAI’s official feed records an October 2 GPT-6 family guide for choosing models, tuning reasoning effort, improving prompts and skills, coordinating tools and preparing startup workflows for production.
  2. OpenAI API model catalogOpenAI Developers · Reviewed: Current featured-model ordering; Astra, GPT-6.1 Sol and Luna positioning; model identifiers; common text-and-image input, text output, multilingual and vision availability; Responses API access · Retrieved · Supports: The live catalog recommends GPT-6 Astra for complex reasoning and coding, GPT-6.1 Sol for balancing intelligence and cost, and GPT-6 Luna for cost-sensitive high-volume work; all three are listed as available through the Responses API.
  3. GPT-6 Luna model referenceOpenAI Developers · Reviewed: Model ID; positioning; reasoning efforts; modalities; context and output limits; knowledge cutoff; Standard token prices; long-context threshold; processing modes; endpoints; tools; EU residency; default rate-limit table · Retrieved · Supports: OpenAI documents gpt-6-luna with a 1,050,000-token context window, 922,000-token maximum input, 128,000-token maximum output, six reasoning settings, Responses API tools, EU residency qualifications and Standard rates of $0.10 input and $0.50 output per million short-context tokens.
  4. OpenAI API pricingOpenAI Developers · Reviewed: Standard, Batch, Flex and Fast GPT-6 family tables; short- and long-context input, cached-input, cache-write and output rates; Astra, GPT-6.1 Sol and Luna comparisons · Retrieved · Supports: The current tables list short-context Standard prices per million tokens of $10 input and $50 output for Astra, $2 and $10 for GPT-6.1 Sol, and $0.10 and $0.50 for Luna; Batch and Flex halve listed Standard rates, while Fast doubles them.
  5. Model selectionOpenAI Developers · Reviewed: GPT-6.1 Sol recommendation; model-and-effort ladder from Luna Low through Astra Extra high; workload frequency, latency, output use and quality questions; same-input comparison guidance · Retrieved · Supports: OpenAI presents its model-and-effort ladder as a starting point, associates Luna with lower-cost bounded work, Sol with complex revisable work and Astra with demanding work, and advises comparing the same inputs while keeping the lightest setting that meets the quality bar.

Limitations

The canonical OpenAI editorial guide returned a Cloudflare JavaScript-and-cookie challenge during review. AccessAllGPT verified its title, description, URL, Product category and October 2 publication timestamp through OpenAI’s official News RSS feed, then reviewed the live developer catalog, Luna reference, pricing tables and model-selection guide directly. We did not call GPT-6 Luna, GPT-6.1 Sol or GPT-6 Astra; verify account access or residency eligibility; reproduce OpenAI evaluations; compare quality, latency, tool behavior, context recall or safety; inspect training data; test rate-limit exhaustion; calculate a reader’s invoice; or conduct a security, privacy, legal or contractual audit. Family positioning is OpenAI guidance, not an independent ranking. Models, snapshots, prices, limits, tools, policies and availability can change.

Disclosures

AccessAllGPT did not receive OpenAI credentials, credits, model access, early information, a briefing, demo, benchmark data, review or compensation for this article. OpenAI did not sponsor, review or endorse it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with OpenAI. Publication-wide relationships are listed on the disclosures page.

Further AccessAllGPT guidance

  1. GPT-6 Astra Ultrafast Is Live at 6× Standard Token Prices
  2. OpenAI Launches GPT-6.1 Sol at $2 Input and $10 Output
  3. GPT-6 Astra: Test the Agent Controls Before You Trust the Benchmark Lead
  4. Choose a Model Without Chasing the Leaderboard
  5. Move an LLM API Without Breaking Production
  6. Design an Agent Benchmark That Predicts Production
  7. AI API Data Retention and Residency: Set the Procurement Gates
  8. AccessAllGPT Research methodology
  9. Publication disclosures

Continue the research

Get evidence-led updates for teams making production AI decisions.