Terra
gpt-5.6-terra
- Input
- $2.00
- Cached
- $0.20
- Output
- $12.00
Free and Go move to Luna with an announced unlimited-text and Think rollout; Plus and Pro get an updated ChatGPT-only Sol with a thought slider. Work, Codex, the API, and Flowith remain separate version or access checks.
API price record · July 30, 2026
These are OpenAI's Standard API rates per 1 million tokens for short-context requests. Prompts above 272K input tokens, other processing modes, cache writes, tools, data residency, and third-party platforms can use different rates.
gpt-5.6-terra
gpt-5.6-luna
API token prices are separate from ChatGPT Work and Codex subscriptions. OpenAI says subscription prices and quota budgets stayed unchanged, while Terra and Luna now consume fewer credits on paid subscriptions.
Family map
The tier names are durable roles. Exact limits, pricing, and features should always be checked in the live OpenAI docs.
gpt-5.6-sol
Complex production work where quality, coding, knowledge work, and difficult multi-step execution matter most.
gpt-5.6-terra
Strong general performance when a workflow needs a lower-cost role than the flagship tier.
gpt-5.6-luna
High-volume or latency-sensitive work where efficiency is the primary model-selection constraint.
Sol processing choice
Fast is the documented premium processing mode for ordinary API requests. Ultrafast is a separate selected-customer Sol API preview powered by Cerebras; its provider-reported output rate is not an end-to-end latency, price, SLA, or access guarantee.
Default choice when cost matters more than premium latency.
No Fast mode opt-in required.
Use as the evaluation baseline for latency and completed-task cost.
High-value, user-facing Sol work where latency is worth a premium.
Set service_tier to "fast"; "priority" remains compatible.
Up to 2.5× faster for Sol at twice the Standard token price; ramp limits can downgrade traffic.
Selected-customer Sol API work where output speed is the defining constraint.
No public request parameter documented in the announcement.
Provider-reported up to 750 output tokens/s and up to 14× Standard; price, SLA, fallback, regions, and broader availability are not stated.
Requests using the legacy priority value continue to work. Fast traffic can be downgraded to Standard when ramp-rate limits apply; inspect the returned service tier and usage data when latency or billing matters.
The Ultrafast announcement does not document a public request value. Do not infer one from the Fast API or present the preview as available in Codex, ChatGPT, Flowith, or third-party providers.
Availability boundary
The August 6 announcement changes ChatGPT access and controls; it does not update Work or Codex Sol, define a new API revision, or establish a Flowith launch, entitlement, free-credit offer, or model mapping.
Luna becomes the default during the August 6 rollout week. Unlimited text and a Think button begin the following week, subject to abuse guardrails; uploads, images, and other tools retain limits.
ChatGPT plan access only; not evidence of Flowith access.
The updated ChatGPT-only Sol powers quick and deeper responses. A slider on web, mobile, and desktop controls how much thought ChatGPT applies.
A ChatGPT control, not an API or Flowith setting.
OpenAI says the Sol version powering Work did not change in the August 6 ChatGPT release.
Separate OpenAI product and version boundary.
OpenAI says the Sol version powering Codex did not change in the August 6 ChatGPT release.
Separate OpenAI product and version boundary.
OpenAI documents Sol, Terra, and Luna model IDs. Ultrafast for Sol is a limited preview for selected API customers, not general API availability.
Does not confirm a Flowith model mapping.
No internal launch evidence was found for GPT-5.6 during this route review.
Verify in the live workspace model selector.
Selection protocol
Start with Sol for frontier capability, Terra for balanced work, or Luna for efficient throughput. Do not collapse a multi-model workflow into one tier without evaluation.
ChatGPT, Codex, the OpenAI API, and Flowith have separate access and entitlement boundaries. Verify the surface you will actually use.
Use Standard as the baseline. Test documented Fast mode when latency is worth the premium. Treat Ultrafast as a separate selected-customer preview until access, request, pricing, and operating terms are documented for your account.
Measure task success, output completeness, latency, token use, tools, and structured-output behavior on representative work before switching.
Primary sources
Use the announcement for the availability event and the model guide for current implementation details.
Related model routes