OpenAI just cut the price of a model it launched three weeks ago — by as much as 80%. On July 30, GPT-5.6 Luna, the fast and cheap tier of the GPT-5.6 lineup, dropped from $1.00 to $0.20 per million input tokens and from $6.00 to $1.20 per million output tokens. GPT-5.6 Terra, the mid-tier model most production apps actually run on, got a smaller cut: $2.50 to $2.00 per million input tokens, $15.00 to $12.00 per million output tokens, a 20% reduction. The flagship Sol model didn't move at all. Instead it got a new 'Fast' tier that costs twice as much for roughly 2.5x the speed — the opposite direction.
Why discount a model you just shipped?
Three weeks is not a long product life before a repricing, and OpenAI didn't frame this as a promotion. In its own developer announcement, the company credited production GPU kernel improvements and a smarter speculative-decoding pipeline — the technical work that makes a model cheaper to actually run, not a marketing budget. CFO Sarah Friar has said customers care about 'the cost of a successful outcome, including the time, retries, oversight, and errors,' not the sticker price on a token. That's a fair point when you're selling API access. It's also the kind of framing a company uses when it doesn't especially want to say the word 'competition' out loud.
My honest read: this is defense, not generosity. OpenAI just crossed a billion active users and two million business customers, and it set these prices the same month Chinese labs keep undercutting everyone on cost per token. Kimi K3 shipped as a 2.8-trillion-parameter open model almost nobody outside a data center can actually run — but its API pricing still drags the whole market down. Alibaba's Qwen3.8-Max is edging toward a full open-weight release. When your closest rivals for developer mindshare give reasoning-grade models away at a fraction of your price, an 80% cut on your cheap tier isn't charity. It's triage.
GPT-5.6 pricing, per million tokens
Luna
Model
$1.00
Old input
$0.20
New input
$6.00
Old output
$1.20
Terra
Model
$2.50
Old input
$2.00
New input
$15.00
Old output
$12.00
Sol (standard)
Model
unchanged
Old input
unchanged
New input
unchanged
Old output
unchanged
Sol Fast (new)
Model
2x standard
Old input
2x standard
New input
2x standard
Old output
2x standard
Model
Old input
New input
Old output
New output
Luna
$1.00
$0.20
$6.00
$1.20
Terra
$2.50
$2.00
$15.00
$12.00
Sol (standard)
unchanged
unchanged
unchanged
unchanged
Sol Fast (new)
2x standard
2x standard
2x standard
2x standard
Sol didn't get cheaper — it got a faster, pricier lane
The detail worth sitting with is what didn't change. Sol is OpenAI's actual flagship, the model you reach for when Luna and Terra aren't smart enough. It kept its price and picked up a Fast mode that costs double for roughly 2.5x the throughput — priced as a premium, not a discount. Cut prices on the commodity tiers where competition is fiercest, hold the line on the model that can't easily be swapped for an open-weight alternative. That split isn't an accident.
OpenAI credits the price cut to real efficiency gains in how it serves GPT-5.6 on its own GPU fleet. · Unsplash
What this changes if you're actually paying for tokens
If you're running a Terra-based product, your inference bill drops 20% with zero code changes. If you're on Luna for high-volume, latency-sensitive tasks — classification, extraction, cheap chat — you're now paying prices that compete directly with open-weight models hosted on someone else's GPUs, not just other closed labs. The change also flows through to ChatGPT and Codex subscriptions: plan prices stay the same, but Terra and Luna usage now consumes fewer credits against your quota, so the same monthly plan effectively stretches further.
One quiet catch: vision token costs went up 20% across the entire GPT-5.6 family in the same update. If your workload leans on image input — screenshots, document scans, anything multimodal — you're not getting the discount the headline promises. Check your own usage pattern before assuming this is a straightforward win.
Quick answers
Did ChatGPT Plus or Pro subscription prices change?
No. OpenAI held consumer subscription pricing steady — the cuts apply to API token pricing and to how much of your subscription quota Luna and Terra usage consumes.
Is GPT-5.6 Sol getting cheaper too?
Not at standard speed. Sol's base price is unchanged; the only new option is a Fast mode priced at double for roughly 2.5x the throughput.
Why did OpenAI cut prices right after launch instead of at launch?
OpenAI says the cuts come from real infrastructure efficiency gains — better GPU kernels and speculative decoding — made after GPT-5.6 had already been running in production for a few weeks.
Does this mean AI is getting cheaper across the board?
For OpenAI's cheaper tiers, yes. But it's happening under real competitive pressure from Chinese open-weight labs, not general industry generosity — expect more moves like this, not a permanent price floor.
This is a prediction, not a report: expect at least one more Luna-tier price cut from a major closed lab before the end of 2026. The floor for 'cheap, fast, good enough' inference keeps dropping because open-weight competition keeps forcing it down, and OpenAI just showed it's willing to reprice a three-week-old model rather than cede that ground. Whether that's good for OpenAI's margins is a separate question from whether it's good for you. For now, it's very good for you.