Terra
gpt-5.6-terra
- Input
- $2.00
- Cached
- $0.20
- Output
- $12.00
OpenAI now prices Standard short-context Terra at $2 input and $12 output per million tokens, and Luna at $0.20 input and $1.20 output. Sol adds Fast mode for premium latency. Flowith availability remains a separate live check.
API price record · July 30, 2026
These are OpenAI's Standard API rates per 1 million tokens for short-context requests. Prompts above 272K input tokens, other processing modes, cache writes, tools, data residency, and third-party platforms can use different rates.
gpt-5.6-terra
gpt-5.6-luna
API token prices are separate from ChatGPT Work and Codex subscriptions. OpenAI says subscription prices and quota budgets stayed unchanged, while Terra and Luna now consume fewer credits on paid subscriptions.
Family map
The tier names are durable roles. Exact limits, pricing, and features should always be checked in the live OpenAI docs.
gpt-5.6-sol
Complex production work where quality, coding, knowledge work, and difficult multi-step execution matter most.
gpt-5.6-terra
Strong general performance when a workflow needs a lower-cost role than the flagship tier.
gpt-5.6-luna
High-volume or latency-sensitive work where efficiency is the primary model-selection constraint.
Sol processing choice
OpenAI says GPT-5.6 Sol can run up to 2.5× faster in Fast mode than Standard processing at twice the token price, with no intelligence change. It is a processing choice, not a different model.
Default choice when cost matters more than premium latency.
No Fast mode opt-in required.
Use as the evaluation baseline for latency and completed-task cost.
High-value, user-facing Sol work where latency is worth a premium.
Set service_tier to "fast"; "priority" remains compatible.
Up to 2.5× faster for Sol at twice the Standard token price; ramp limits can downgrade traffic.
Requests using the legacy priority value continue to work. Fast traffic can be downgraded to Standard when ramp-rate limits apply; inspect the returned service tier and usage data when latency or billing matters.
Availability boundary
OpenAI's announcement establishes availability on its own surfaces. It does not establish a Flowith launch, account entitlement, free-credit offer, or model mapping.
OpenAI says paid-plan Terra and Luna usage now consumes fewer credits; subscription price and quota budgets did not change.
Not evidence of Flowith access.
Terra remains available to Free and Go; paid plans can choose Terra and Luna, with lower credit consumption for both tiers.
Separate OpenAI product surface.
OpenAI documents Sol, Terra, and Luna model IDs for API use.
Does not confirm a Flowith model mapping.
No internal launch evidence was found for GPT-5.6 during this route review.
Verify in the live workspace model selector.
Selection protocol
Start with Sol for frontier capability, Terra for balanced work, or Luna for efficient throughput. Do not collapse a multi-model workflow into one tier without evaluation.
ChatGPT, Codex, the OpenAI API, and Flowith have separate access and entitlement boundaries. Verify the surface you will actually use.
Use Standard as the baseline. Test Fast mode for latency-sensitive Sol requests, then verify the response service tier and billed rate instead of assuming every request received premium processing.
Measure task success, output completeness, latency, token use, tools, and structured-output behavior on representative work before switching.
Primary sources
Use the announcement for the availability event and the model guide for current implementation details.
Related model routes