ALITEQ.

OpenAI just cut its AI prices by 80% three weeks after launch and China is the reason

GPT-5.6 Luna dropped from $1 to 20 cents per million tokens overnight. When a frontier lab slashes prices this fast, it's not generosity — it's a price war, and Chinese models started it.

Lena FischerUpdated 2h ago9 min read
An abstract representation of an AI processor

When a frontier AI lab cuts its prices by 80% just three weeks after a launch, that's not a routine discount — it's a flinch. On July 30, OpenAI slashed the price of GPT-5.6 Luna from $1/$6 to $0.20/$1.20 per million input/output tokens, an 80% cut on the input side, while trimming the mid-tier Terra 20% and leaving the flagship Sol untouched. GPT-5.6 only launched on July 9. Labs don't gut the price of a three-week-old model out of confidence. They do it because someone cheaper is eating their lunch — and in this case, that someone is China.

Why OpenAI blinked

OpenAI's official line is that efficiency improvements — models and infrastructure getting cheaper to run — made the cut possible. That's true and beside the point. The forcing function is competition, specifically from China. Chinese models have reportedly captured around 46% of US enterprise token usage on OpenRouter, at times peaking above US-origin models, and prices like DeepSeek V4 Pro at $0.435/$0.87 per million tokens (helped by a standing 75% promotional discount) make premium US pricing hard to justify for high-volume work. Enterprises, meanwhile, have gotten cost-sensitive — less willing to deploy expensive models without a clear return. Put those together and you get exactly this: a frontier lab defending market share by racing down on price, three weeks after launch. The pricing power that frontier labs enjoyed is eroding fast, and this cut is the clearest proof yet.

A data center corridor with servers
An 80% cut three weeks after launch isn't generosity — it's a price war, and cheap Chinese models started it. · Unsplash

What it means for you — and for local AI

There are two takes here, and I hold both. For anyone building on cloud AI, this is great news — costs are falling fast, and high-volume work that was uneconomical is suddenly viable. But here's the counterintuitive part I think matters most: cheaper cloud AI strengthens the case for running AI locally, it doesn't kill it. Why? Because the price war is a race to the bottom on a metered service, and metered services can raise prices again once the competition shakes out — whereas a local model on your own hardware is a fixed cost you control, private, and immune to whatever the API market does next. The open models driving this price war (DeepSeek and others) are the same models you can run yourself. So the honest framing is: cloud AI is getting cheap, which is good, but the reason it's getting cheap — powerful cheap open models — is exactly what makes local AI more viable too. If you run the numbers, high-volume users still often win by owning the hardware.

Quick answers

How much did OpenAI cut GPT-5.6 prices?
On July 30, 2026, OpenAI cut GPT-5.6 Luna — its fastest, most cost-effective tier — by 80% on input, from $1 to $0.20 per million input tokens, and from $6 to $1.20 per million output tokens. The mid-tier Terra was cut about 20% (from $2.50/$15 to $2/$12), while the flagship Sol was left unchanged at $5/$30 per million tokens. The cuts came just three weeks after the GPT-5.6 family launched on July 9.
Why did OpenAI cut its AI prices?
Officially, OpenAI cited efficiency improvements in its models and infrastructure. The real driver is competition, especially from cheap Chinese models, which have captured roughly 46% of US enterprise token usage on OpenRouter and undercut US pricing dramatically — DeepSeek V4 Pro, for example, runs at $0.435/$0.87 per million tokens with a 75% discount. Enterprises have also grown cost-sensitive. Cutting a three-week-old model's price by 80% is a defensive move to protect market share in an intensifying AI price war.
Does cheaper cloud AI make local AI pointless?
No — arguably the opposite. Cheaper cloud AI is driven by powerful, inexpensive open models (like DeepSeek), and those are the same models you can run locally on your own hardware. Cloud pricing is a metered service that can rise again once competition settles, while a local model is a fixed cost you control, private, and independent of the API market. For high-volume users especially, running AI on your own hardware often still wins on total cost, and it adds privacy that no cloud service offers.

OpenAI cutting prices 80% three weeks post-launch is the loudest signal yet that the AI price war is real and China is driving it. Cheaper cloud AI is good — but it makes running open models locally more compelling, not less, since it's the same cheap open models underneath. Weigh it with our local vs cloud cost breakdown. Sources: CNBC and VentureBeat.

AI & Local Compute Editor

Lena Fischer

Lena runs more GPUs at home than she'll admit to and has quantized more models than she's finished reading about. She writes about running AI on your own hardware — what actually fits, what's genuinely fast, and what the polished cloud demos quietly leave out.

Work out the hardware

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading