GPT-5.6 Luna

OpenAIfast New-80%

The fastest, cheapest member of the GPT-5.6 family and the site of the steepest price cut of the year from a US lab: on 30 Jul 2026 OpenAI dropped Luna 80%, from $1/$6 to $0.20 input / $1.20 output per 1M tokens. Cache reads run $0.02. That is an explicit answer to DeepSeek and the Chinese open-weight tier on price — though at $0.20/$1.20 Luna is still comfortably more expensive than DeepSeek V4 Flash at $0.14/$0.28. Full 1.05M context, unlike most rivals' cheap tiers.

Context Window

1M

tokens

Max Output

128K

tokens

Input Price

$0.2

per 1M tokens

Output Price

$1.2

per 1M tokens

Speed

186

tokens/sec

Released

Jul 2026

2026-07-09

Blended Cost

$0.70

per 1M tokens

Value Score

118.6

quality per $

Capabilities

ChatVisionFunction CallingCode Generation

Benchmarks

Quality Index
83
MMLU Pro
84.2
HumanEval (Coding)
87.6
MATH
82
Arena ELO
1436

Pricing History

Jul 9, 2026$1 / $6(launch)
Jul 30, 2026$0.2 / $1.2current

About GPT-5.6 Luna

GPT-5.6 Luna is a fast AI model by OpenAI, released on July 9, 2026. It supports a context window of 1050K tokens and can generate up to 128K output tokens.

At $0.2 per million input tokens and $1.2 per million output tokens, its blended cost of $0.70/1M tokens makes it one of the most affordable models available. Its value score of 118.6 reflects the balance of quality and cost.

Using GPT-5.6 Luna with Swfte

Access GPT-5.6 Luna through Swfte Connect, our unified LLM gateway. Connect gives you a single API for 50+ models, with automatic routing, cost optimization, and fallback handling. You can also try GPT-5.6 Luna in our AI Playground before integrating.