Model detail
GPT-5.6 Terra
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic tasks where capability and cost need to be balanced, offering strong performance at roughly half the cost of Sol.
Model specs
- Context length
- 1.05M
- Max output
- 128K
- I/O modalities
- Text / Image
- Released
- 2026-07
02
API endpoints
One MixRoute gateway, OpenAI-compatible
-
OpenAI-compatible
/v1/chat/completionsPOST
03
Tiered pricing
Unit: /1M Tokens
| Tier | Input /1M Tokens | Output /1M Tokens | Cache read /1M Tokens | Cache write /1M Tokens |
|---|---|---|---|---|
| standard Length ≤ 272K | $2.0000 | $12.0000 | $0.2000 | $2.5000 |
| long_context Length > 272K | $4.0000 | $18.0000 | $0.4000 | $5.0000 |
05
Selection summary
Quickly judge whether it fits your workload
What is this model good for?
Everyday coding, reasoning, and agent tasks to test first
Use GPT-5.6 Terra for everyday coding, reasoning, and agent workflows where capability and cost need to be balanced, as stated in the official description. The full specifications and capability list are not published yet, so run one representative coding or reasoning task and compare output quality against your current model before scaling.
model: "gpt-5.6-terra", then compare with your baseline results.
What should you check before using it?
Check the 272K pricing tier boundary
Before adopting GPT-5.6 Terra, check the pricing boundary at 272K input tokens: requests at or below that length are billed $2.0000 per 1M input tokens, while longer requests move to the long-context tier at $4.0000. Measure your typical request size and estimate both tiers so the projected cost reflects your real usage.
Why use it through MixRoute?
Call it on the OpenAI-compatible endpoint with its model ID
Call GPT-5.6 Terra through MixRoute's OpenAI-compatible endpoint (POST /v1/chat/completions) with the model ID gpt-5.6-terra. Because the endpoint follows the OpenAI request format, you can keep your existing OpenAI-style client, change the model string, and test the same code path against your current setup.
/v1/chat/completions with model: "gpt-5.6-terra" and rerun your integration tests.
08
Token cost estimator
Live estimate from this page's pricing, not an actual bill
Share of the same prefix read repeatedly, up to 100%
09
FAQ
For requests at or below 272K tokens, GPT-5.6 Terra charges $0.2000 per 1M tokens for cached reads instead of $2.0000 for standard input; above 272K, the long-context tier charges $0.4000 cache read versus $4.0000 input. The cache-read price is one tenth of the input price in both tiers. Measure how many input tokens repeat across your requests, then estimate with your own cache-to-input ratio before assuming a budget change.
One endpoint, a testable decision
Test this model and alternate routes with the same request format
Start from a real workload, then let quality, total cost, and failure conditions decide whether to send production traffic.