If you build products on top of frontier models, you have probably been waiting for the same thing I was: prices falling steadily as competition arrives. It is what happened to bandwidth, to storage, to compute. It seemed like a matter of time.
It has not really happened, at least not in the way we expected. The reason is visible in the price lists, if you stop reading the model names and start reading the numbers.
Two things follow from those numbers, and this piece is about both. The first is that prices are not really falling, they are standing still while what you get for them moves. The second is what that implies about where the labs are actually competing, which is no longer the model.
Four numbers, three generations
Here is OpenAI's public price list, sorted by input price, in US dollars per million tokens.
| Model | Input | Output |
|---|---|---|
| gpt-5.4-nano | $0.20 | $1.25 |
| gpt-5.6-luna | $0.20 | $1.20 |
| gpt-5.4-mini | $0.75 | $4.50 |
| gpt-5.6-terra | $2.00 | $12.00 |
| gpt-5.4 | $2.50 | $15.00 |
| gpt-5.5 | $5.00 | $30.00 |
| gpt-5.6-sol | $5.00 | $30.00 |
| gpt-5.4-pro | $30.00 | $180.00 |
| gpt-5.5-pro | $30.00 | $180.00 |
Four prices. $0.20, about $2, $5, and $30. Three generations of models keep landing on them, with one straggler: gpt-5.4-mini at $0.75 sits between rungs, and the $2 rung is really $2.00 to $2.50. The pattern is a pattern, not a law.
The flagship has been $5 / $30 since GPT-5.5. The pro tier has been $30 / $180 since GPT-5.4. The cheapest rung has been $0.20 since nano.
The prices are fixed rungs on a ladder. What changes is how much intelligence stands on each one.
So what was the "80 percent price cut"?
On 30 July 2026 OpenAI cut GPT-5.6 Luna by 80 percent, from $1 / $6 to $0.20 / $1.20, twenty-one days after the family launched. Terra fell 20 percent. Sol, the flagship, did not move.
Read against the ladder, the headline turns into something more useful. OpenAI did not invent a cheaper price. It moved a capable model down onto a rung that already existed, the one nano occupied a generation earlier, at a nearly identical number.
Which tells you what to measure. Nano and Luna cost the same. The question is not what changed about the price. It is what you now get for it.
Prices did not simply fall
One vendor's price list could be a quirk of how OpenAI packages models. It is not. Widen the view to every major lab and the same shape appears, with a wrinkle.
The frontier fell once, hard, then stopped. Anthropic's Opus went from $15 / $75 at 4.1 to $5 / $25 at 4.5, a 67 percent cut, and has stayed at $5 / $25 across five consecutive versions, 4.5 through 5. OpenAI's flagship has held $5 / $30 across two generations.
Meanwhile several things went up. Claude Haiku went from $0.80 / $4 at 3.5 to $1 / $5 at 4.5. And Moonshot's Kimi K3 costs $3 / $15 against K2.6 at $0.95 / $4, one generation apart, from the lab most associated with cheap Chinese models.
A third group is cheap in a way that comes with an expiry date. Claude Sonnet 5 runs $2 / $10, introductory pricing published as ending 31 August 2026, after which standard $3 / $15 applies. Alibaba lists Qwen3.7 Max at $2.50 / $7.50 on its Singapore endpoint with a limited-time 50 percent discount applied on top. MiniMax M3 is $0.30 / $1.20, which its own docs describe as a permanent 50 percent cut from the original rate. Read a discounted number as a list price and you will plan a budget around a rate the vendor has already told you is temporary.
So nominal prices are flat to rising, while intelligence per dollar collapses. Both are true at once, and conflating them is why "prices will come down" has felt wrong for a year.
What each rung buys you now
If the rungs hold still and only capability moves, then the only question worth asking about a price is what currently stands on it. That needs measurement rather than list prices, so here is the same market plotted against what the models can actually do.
Every point is Artificial Analysis data read on 30 July 2026: their Intelligence Index score against their measured cost to complete one Index task. The teal line is the Pareto frontier, the set of models where nothing else is both smarter and cheaper. Grey points are dominated, meaning you can buy something better for less.
On the $0.20 rung, Luna now scores 51 and finishes an Index task for six cents. Its peers at that score:
| Model | Index | Cost per task |
|---|---|---|
| GPT-5.6 Luna (max) | 51 | $0.06 |
| GLM-5.2 (max) | 51 | $0.27 |
| Claude Opus 5 (low) | 51 | $0.36 |
Four and a half times cheaper than the next model at the same measured intelligence. The rung stayed still; what stands on it did not.
Ask that same question at every level of the Index, not just the one Luna landed on, and the ladder reappears in a different form. Here is the cheapest way to buy each level of measured intelligence:
| Index | Cheapest model at that level | Cost per task |
|---|---|---|
| 38 | GPT-5.6 Luna (medium) | $0.01 |
| 40 | DeepSeek V4 Flash | $0.02 |
| 44 | DeepSeek V4 Pro | $0.04 |
| 51 | GPT-5.6 Luna (max) | $0.06 |
| 54 | Grok 4.5 (high) | $0.35 |
| 56 | Claude Opus 5 (medium) | $0.62 |
| 57 | Kimi K3 | $0.72 |
| 59 | Claude Opus 5 (high) | $1.06 |
| 60 | Claude Opus 5 (xhigh) | $1.56 |
| 61 | Claude Opus 5 (max) | $2.03 |
The rungs are stable and the occupants are not. Anthropic holds every level from 56 upward except one, where Kimi K3 sits. OpenAI holds the two ends of the cheap range, 38 and 51, and DeepSeek holds the two rungs in between. Luna took 51 this month.
Two things fall out, and both are about occupancy rather than price. GPT-5.6 Sol scores 59 at $1.54 a task, while Claude Opus 5 on high scores the same 59 for $1.06. Sol costs 45 percent more for the same measured intelligence. And because Luna now sits at 51 for six cents, it dominates Gemini 3.6 Flash, Gemini 3.1 Pro Preview, GLM-5.2, Claude Opus 5 on low, and Qwen3.7 Max. Five models across four labs became hard to justify in an afternoon, and none of their vendors changed a price.
Nobody is competing at the top
That table also says something about where the fighting is. The cheap rungs are crowded, contested and changing hands in an afternoon. The top of the ladder is not.
Notice what did not happen this month. Sol is $5 / $30, exactly GPT-5.5's price. Fable 5, Opus 5 and the $30 / $180 pro tiers went untouched. Every cut landed below the flagship: 80 percent at the bottom, 20 percent in the middle, zero at the top.
Price competition is evidence that buyers have a choice. At the bottom of the ladder they plainly do. At the top, apparently, they do not. Which raises the obvious question: if the labs are not fighting each other on the price of their best models, what are they fighting over?
The part that matters if you build on this
They are fighting over the harness: the scaffolding that connects a model to tools, context, memory and a real task. OpenAI say so themselves. Their efficiency gains come from "the models, the inference systems that run them, and the agentic harness that connects them to tools and context."
This is where the ladder stops being a pricing observation and starts being your problem, because a lab that owns the model has a structural advantage in harness that has nothing to do with being better at building one.
Three facts from the price lists rather than from theory:
- The cut reached their own harness first. OpenAI: the new prices are "also reflected in how usage is counted against paid subscriptions when using Codex and ChatGPT Work." Their harness receives it as credits. You receive a list price.
- There is a harness-specific model on the price list.
gpt-5.3-codexat $1.75 / $14, a SKU shaped for their own product. - Anthropic bills harness time as its own axis. Claude Managed Agents charge $0.08 per session-hour on top of tokens. A third-party harness cannot recreate that line item, because it pays tokens and runs its own infrastructure.
For a lab, tokens never cross a market. Training and serving are cost centres, and the harness is the product. For everyone else the API list price is the input cost, and it is set by the company you compete with on harness. Your cost of goods is their transfer price.
That is the real reason building on frontier models has not felt cheaper even as the models got much better. The rung you buy at is flat. The rung they buy at is internal.
On subsidies, honestly
There is a version of that argument that goes further than the evidence does, so here is the limit of it. Saying a lab's internal cost is lower than your list price is not the same as saying the list price is below cost, and I cannot show you the second one. No lab publishes per-token gross margin, so nobody outside these companies knows.
Two things are documented and worth keeping separate. State money is in this market: DeepSeek's 2026 raise was backed by state-linked investors, including affiliates of the third phase of the China Integrated Circuit Industry Investment Fund. That is equity in a company, which is not the same claim as a subsidy on a token, and no published figure connects the two.
The same gap shows up in OpenAI's own numbers. They stated a 20 percent reduction in serving cost from kernel optimisation work, then cut Luna by 80 percent. Those figures do not reconcile, and they are not even about the same model. The difference is equally consistent with margin compression, cross-subsidy from subscriptions, or simply buying the low end of the market. From outside they are indistinguishable, and anyone who tells you which one it is, is guessing.
Four ways this comparison goes wrong
Everything above rests on comparing prices to each other, which is easier to get wrong than it looks. Four mistakes are common enough to name, and each one flips the conclusion.
Comparing across capability tiers. Kimi K3 at $3 beside Luna at $0.20 makes Chinese models look expensive. They are six index points apart and are not alternatives to each other. At 57 for $0.72 a task, K3 sits on the frontier. Always ask what else scores the same.
Comparing across rungs inside one lab. "OpenAI cut prices" is true and useless. The bottom fell and the top did not move. Which rung moved is the whole story.
Treating the token as a unit. Anthropic's own docs state that Claude 4.7 and later use a newer tokenizer producing "approximately 30 percent more tokens for the same text" than earlier Claude models. A price per million tokens is only comparable between two models that cut text the same way, and that baseline is Claude against earlier Claude, not Claude against OpenAI. Cost per completed task sidesteps the unit entirely, which is why this piece uses it.
Reading a list price as what you will pay. Above 272,000 input tokens OpenAI's input and cached-input prices double and output rises by half, so Luna becomes $0.40 / $1.80. Not an OpenAI quirk: xAI doubles at 200k, Gemini 3.1 Pro doubles at 200k, and MiniMax doubles at 512k, to $0.60 / $2.40. Anthropic is the exception worth knowing, billing the full 1M window at standard rates from Claude 4.6 onward. A rung is a starting point, and on most vendors long-context work quietly puts you on the next one up.
What to do with this
If the rungs are fixed and only capability moves, the useful habit is not watching prices. It is periodically asking one question:
At the rung I am already paying for, what is the best thing standing on it today?
That question would have caught this week's change without reading a single announcement, and it will catch the next one. It also surfaces the more expensive and more common mistake: a model you deployed six months ago that quietly became the wrong choice while its price stayed exactly the same.
The second half of this is harder to act on, and more important. If the rung you buy at is fixed and the rung your competitor buys at is internal, then no amount of shopping around fixes the gap. What you build on top of the model has to be worth more than the discount you will never get.
Method, and what I left out
Vendor list prices came from each vendor's own pricing page on 30 July 2026: OpenAI, Anthropic, Google, xAI, DeepSeek, Z.ai, Moonshot, Alibaba and MiniMax. No aggregator was used as a source for any figure. Where a vendor flags a rate as promotional or introductory, the list price is quoted and the discount noted.
Intelligence Index scores and cost per task are Artificial Analysis measurements read the same day. Their Luna page states $0.20 / $1.20, so the cost figures reflect post-cut prices as measured, not launch-day numbers rescaled by hand. Rescaling would have been wrong: Artificial Analysis re-ran the evaluation rather than adjusting arithmetic, and a re-run can move the score as well as the cost.
Cost per Index task is a benchmark workload, not yours. It carries a particular input-to-output ratio and a particular reasoning effort, and a workload dominated by long cached prompts will rank these models differently. It is a much better comparator than dollars per million tokens, and it is still a proxy. Evaluate on your own traffic before moving anything in production.
Every chart and diagram in this post is drawn in code from the numbers above. Only the header illustration is generated, because a generated chart in an article about prices would be a made-up chart.
Short, source-checked breakdowns like this go out most days as sojho — @sojho.dev on Instagram, TikTok, X, LinkedIn and YouTube. Same rule there as here: primary sources only, and the workings are published.