Cloud & Software

OpenAI Cuts GPT-5.6 Prices Up to 80% — but Not Where You Think

Luna drops 80%, Terra 20%, and flagship Sol holds its price while gaining a double-cost Fast Mode. The real story is segmentation: the commodity tier is a knife fight, the frontier is not for sale cheap.

· Jul 31, 2026
OpenAI Cuts GPT-5.6 Prices Up to 80% — but Not Where You Think
Illustration generated by AI
Table of contents
  1. The new price list
  2. How OpenAI says it paid for this
  3. The strategy under the numbers
  4. What it means for buyers

OpenAI cut API prices across its GPT-5.6 lineup on July 30 — three weeks after the models went generally available. The headline number is dramatic: up to 80% off. The detail that matters more is which models got it.

The new price list

Model Old (in/out per 1M tokens) New (in/out) Change
GPT-5.6 Luna $1 / $6 $0.20 / $1.20 −80%
GPT-5.6 Terra $2.50 / $15 $2 / $12 −20%
GPT-5.6 Sol $5 / $30 $5 / $30 unchanged

Cached input reads get a 90% discount ($0.02 per million on Luna, $0.20 on Terra). And Sol — while holding its base price — gained a new Fast Mode, replacing the old Priority Processing: up to 2.5× standard speed at exactly double the price ($10/$60 for Sol, with Fast variants for Terra and Luna too), per Unite.AI's breakdown.

How OpenAI says it paid for this

The efficiency claims behind the cut are specific: kernel-level optimization work that reduced end-to-end serving costs by 20%, plus speculative-decoding improvements that raised token-generation efficiency by more than 15%. The detail several outlets have seized on: OpenAI credits part of the kernel rewrite to GPT-5.6 Sol itself, making this arguably the first price cut partially engineered by the model being sold. Treat that framing with appropriate salt — "the model funded its own discount" is a great marketing line precisely because it's unverifiable from outside — but the 20% serving-cost figure is the company's own, and the price cut backing it is real.

The strategy under the numbers

Read as a price sheet, this is a discount. Read as a strategy, it's segmentation:

  • The commodity tier is a knife fight. At $0.20 input, Luna undercuts Anthropic's Haiku 4.5 by roughly 5× on input and 4× on output. High-volume, low-stakes workloads — classification, extraction, summarization at scale — are being priced like infrastructure, because that's the tier where switching costs are lowest and volume is biggest.
  • The flagship holds its price and adds a premium above itself. Sol didn't get cheaper; it got a 2×-price express lane. That's classic price discrimination upward: latency-sensitive customers (agents, interactive products) self-identify and pay double for the same tokens, faster.
  • The middle got a courtesy trim. Terra's 20% keeps it plausible against mid-tier rivals without disturbing the structure.

This is the same shape Anthropic used when it halved Opus pricing at launch: aggressive at the volume end, disciplined at the frontier end. Nobody is discounting their best model — they're discounting the models that stop you from leaving.

What it means for buyers

For teams already running GPT-5.6 in production workflows, the practical move is routing, not switching. An 80% cut on Luna changes the math on every workload currently over-served by Terra or Sol — the cost lever in 2026 is matching each task to the cheapest tier that clears your quality bar, then letting caching discounts compound on top. Teams that route aggressively will see their effective per-token spend fall much further than any single list price suggests.

It's also a reminder of the environment these cuts are answering. AI spending is squeezing enterprise IT budgets, and token bills have become visible enough to make headlines — the US Army literally running out of AI tokens was this spring's warning shot. Vendors know procurement is watching unit costs now. An 80% cut on the cheap tier is as much a retention message to CFOs as it is a technical achievement.

The frontier still costs $30 per million tokens out. But the floor just dropped through the floor — and for most of what enterprises actually run through these APIs, the floor is where the volume lives.