GPT-5.6 Luna
After an 80% price cut in July 2026, Luna became the value outlier among frontier-family models: it outperforms the previous generation's top-end Opus tier on coding evals while costing about a fiftieth of Fable 5 per input token. This is the model to route bulk traffic through in a tiered architecture.
Pricing checked against provider documentation on . How we verify
gpt-5.6-luna reasoning Input
$0.20
/ 1M tokens
Output
$1.20
/ 1M tokens
Context
1.05M
1,048,576 tokens
Max output
128K
tokens per request
Pricing caveat: Cut 80% from the $1/$6 launch price on 2026-07-30.
What it costs in practice
A workload of 10M input and 2M output tokens per month — roughly a moderately busy production assistant — runs $4.40/month on GPT-5.6 Luna. Routing that same volume through the batch API brings it to $2.20. The next cheaper option, DeepSeek V4-Flash , would cost $3.52.
Where GPT-5.6 Luna fits
- Classification and content moderation
- High-frequency agent subtasks
- Bulk summarization
- Cost-sensitive chat features
Specifications
- API model ID
- gpt-5.6-luna
- Provider
- OpenAI
- Model family
- GPT-5.6
- Released
- July 2026
- Knowledge cutoff
- February 2026
- Context window
- 1,048,576 tokens
- Max output
- 128,000 tokens
- Native reasoning
- Yes
- Relative latency
- low
- Input modalities
- text, image, audio
- Output modalities
- text
- Open weights
- No
- Cached input
- $0.02 / 1M tokens
- Batch discount
- 50%
Frequently asked
How much does GPT-5.6 Luna cost?
GPT-5.6 Luna costs $0.20 per million input tokens and $1.20 per million output tokens, with cached input at $0.02. Batch processing is 50% cheaper.
What is GPT-5.6 Luna's context window?
GPT-5.6 Luna accepts up to 1,048,576 tokens of context and can generate up to 128,000 output tokens per request.
Is GPT-5.6 Luna a reasoning model?
Yes. GPT-5.6 Luna performs native chain-of-thought before answering. Those thinking tokens are billed at the output rate, so budget above the sticker price.
What is GPT-5.6 Luna best for?
High-volume work on a tight budget. After an 80% price cut in July 2026, Luna became the value outlier among frontier-family models: it outperforms the previous generation's top-end Opus tier on coding evals while costing about a fiftieth of Fable 5 per input token. This is the model to route bulk traffic through in a tiered architecture.