Fourth of five on what the 2026 evidence says once you read past the announcement. Part three found the slowdown the labs asked for was already being enforced by memory suppliers and a grid operator in Texas. This one follows the money instead.
Also in the series: the coding productivity data, the agent containment failures, and what to check before you run an open weight model.
On September 14, Anthropic raised Claude Code’s weekly usage limits by 25 percent. Measured against the limits people had been working under the day before, the same change was a 17 percent cut, because a temporary boost expired on that date. Anthropic published both numbers.
Neither is wrong. They count from different baselines, and the price of the subscription never moved.
I went looking for the AI token price increases everyone expects once the subsidies stop. Several have already happened. Most of them never touched a price list.
Two things are true here at the same time, and nearly every piece written on this subject holds one and drops the other. The cost of buying a fixed amount of AI capability is falling about as fast as anything in computing ever has. The monthly bill for a team doing real work is climbing.
The price list came down, and that part is real
Ramp builds an index from corporate card transactions, which means it measures what companies paid rather than what vendors published. In early September the blended rate across its panel was $0.68 per million tokens, down from $1.15 in March. That is a 41 percent drop in six months, on invoices that cleared.
The capability-adjusted picture is steeper. Four MIT researchers measured the price of reaching a fixed benchmark score in a paper titled “The Price of Progress” and put the decline at “around 5x to 10x per year.”
Anyone arguing that AI is about to cost more has to concede that first, so I concede it. The deflation shows up in cleared card transactions rather than in vendor marketing, which is about as good as evidence gets on this subject.
Then read the next clause of the same abstract.
So why did the bill go up
“The price of running frontier models is rising between 3x to 18x per year due to bigger models and larger reasoning demands.”
One paper, one abstract, two measurements. Cost of a fixed capability falls 5 to 10 times a year. Cost of running the best available model rises 3 to 18 times a year. Your invoice sits in the space between them, and which number you feel depends entirely on whether you keep using last year’s model.
Token volume is the mechanism. A ChatGPT exchange in 2023 ran somewhere between 500 and 2,000 tokens. SemiAnalysis puts a 2026 agentic coding task at roughly 96,000 tokens consumed before the model produces its first line of output. Halve the price and multiply the count by fifty, and the arithmetic goes where you would expect.
Ramp’s own data shows buyers reacting to this. Frontier models fell from 53 percent of token share in August to 45 percent in September, which Ramp attributes to companies setting organization-wide defaults that route work to cheaper models. The cooling effect on frontier usage that everyone expects to arrive with a price rise has instead arrived through procurement policy, eight points in a month.
The unit of billing stopped matching the unit of work somewhere around the point agents started calling themselves in loops. Everyone quotes dollars per million tokens. Nobody buys tokens, they buy finished tasks.

The subsidy nobody put on the invoice
Flat-rate subscriptions are where the real money is being given away, and the labs have said so.
Sam Altman, January 5, 2025, on the $200 ChatGPT Pro plan: “I personally chose the price, and thought we would make some money.” OpenAI was losing money on it because people used it more than he expected.
SemiAnalysis later put a number on the gap. A Pro subscription pushed to its limits can consume compute worth up to $14,000 a month at API rates, against the $200 the subscriber pays. Claude Max lands around $8,000. Break-even on the OpenAI plan sits near 6 percent utilization, which means the light users are paying for the heavy ones and the whole model depends on most people not using what they bought.
Nobody has raised those subscription prices. What they have done instead reads as a list of dates.
Cursor moved from 500 fast requests to a $20 credit pool billed at API rates in June 2025, citing the cost of newer models by name, and apologized a month later. Anthropic added weekly limits on top of five-hour limits in July 2025. Windsurf swapped credits for usage quotas in March 2026. OpenAI restored a five-hour cap on Codex inside ChatGPT Plus in August 2026, saying it “helps smoothen the load on our compute.” Then Anthropic’s 25 percent increase that was a 17 percent cut.
GitHub was the most direct about why. Announcing the move from premium requests to token-metered credits in April 2026, it wrote that “a quick chat question and a multi-hour autonomous coding session can cost the user the same amount. GitHub has absorbed much of the escalating inference cost behind that usage, but the current premium request model is no longer sustainable.” A developer in the comments summarized the deal on offer: “You will get less, but pay the same price.”
Sticker prices held at $10, $20, $100 and $200 across all of it. What the sticker bought shrank on a roughly quarterly schedule.
The AI token price increases already happened
Set the subscriptions aside. On the metered API, where the list price is supposed to only ever fall, three increases are already on the record.
OpenAI launched GPT-5 in August 2025 at $1.25 per million input tokens and $10 per million output. Its current flagship, GPT-5.6 Sol, sits at $5 and $30. That is four times the input price and three times the output price, and the trade press called it OpenAI’s first price rise since 2023.
Google’s Gemini 3 Pro Preview arrived in November 2025 at $2 and $12, against Gemini 2.5 Pro’s $1.25 and $10. Sixty percent up on input. Gemini 3.1 Pro held the new number through July 2026, so it was not a preview-period artifact.
Anthropic ran Claude 3 Opus at $15 and $75 in March 2024 and launched Claude Opus 4 at exactly $15 and $75 thirteen months later. The flagship tier did not get cheaper until Opus 4.5 in November 2025.
One layer down, AWS raised the price of EC2 Capacity Blocks for ML by about 15 percent in January 2026, taking the p5e.48xlarge from $34.61 to $39.80 an hour and citing “supply and demand patterns.” Amazon spent fifteen years building a reputation on never doing that.

The race to the bottom, and the version of it that survives
The suspicion I hear most often from people who run infrastructure goes like this. The labs are selling below cost to take share, the weak ones die, and once three companies are left the prices go up. Anyone who lived through rideshare pricing has seen that movie, and in its strict form the evidence here does not support it.
Margins are the problem. Altman said in August 2025 that “we’re profitable on inference. If we didn’t pay for training, we’d be a very profitable company.” OpenAI’s reported compute margin roughly doubled from 35 percent in early 2024 to 70 percent by October 2025, during the same period list prices were falling. An independent model of serving costs published in August 2025 put API markups between 11.8 and 20.3 times. Losses at these companies come from training runs and data centers, not from the marginal token.
The law is worse for the theory. A predatory pricing claim in the United States has to clear Brooke Group, which requires pricing below cost plus a dangerous probability of recouping the losses later. No regulator has opened a case on those grounds. The FTC’s 2024 inquiry, the CMA’s foundation model review and the EU’s interest in Microsoft and OpenAI all target partnership structure and market power, not price.
There is also a technical problem with the test itself, and it is the most interesting document I found while researching this. The International Competition Network’s own guidance for regulators says long-run average incremental cost is the better measure than average variable cost “when the alleged predatory conduct involves products that have large fixed costs and low marginal costs of production, as in the telecommunication, pharmaceutical, or software industries.” Inference has exactly that shape. The standard screen for predatory pricing goes blind in the cost structure AI happens to have.
So the strict version fails on the numbers. What replaces it is narrower, better documented, and points at the same outcome by a different route.
The intent is documented. In June 2026 the Wall Street Journal reported OpenAI exploring drastic token price cuts “to defend its enterprise turf” against a soon-to-be-public Anthropic. In July, Altman posted that his flagship was “half the price and ~twice as token efficient” as Anthropic’s and that he would be “happy to deliver at one-quarter of the price.” Google cut its AI Plus consumer tier from $7.99 to $4.99 in June 2026, and the coverage said it undercut OpenAI. Anthropic then cancelled a planned billing overhaul that same month, reportedly because moving to metered agent pricing while OpenAI was threatening cuts would have been a bad trade.
And the market is concentrating. Menlo Ventures measures enterprise LLM API spend, and its December 2025 report puts OpenAI, Anthropic and Google at 88 percent of it, up from 69 percent in 2023. Open-weight models fell from 19 percent of enterprise usage to 11 percent over the same period. Chinese open models, which are supposed to be the price floor holding everyone honest, account for about 1 percent of enterprise API usage.
Recoupment is the half of the predatory pricing test that fails in a fragmented market, and this market is not fragmenting. It is doing the opposite, at nineteen points in two years.

What the labs report when you read the caveats
On September 14, the same day those Claude Code limits took effect, Anthropic told investors its gross margins run above 80 percent. The qualifier arrived in the same sentence: before revenue shared with distribution partners including Amazon, and before the cost of training its models.
Those two exclusions are most of the business. Epoch AI looked at whether a single model vintage pays for itself and estimated roughly $5 billion of R&D spent on GPT-5 before launch against roughly $2 billion of gross profit across the model’s commercial life. OpenAI’s own fiscal 2025 numbers show $13.07 billion of revenue against $7.5 billion of cost of revenue, a blended margin near 43 percent once free-tier serving is in the denominator.
A version of the Amodei accounting claim has been going around, that he treats revenue as profit and ignores cost of goods sold. I chased it and it does not hold. His per-model payback argument, made on a podcast in August 2025, explicitly accounts for inference cost. The phrase people are quoting traces back to commentary on Anthropic’s adjusted operating income, which excludes training and partner revenue share but does include the cost of serving.
The honest criticism is narrower and harder to answer. Every margin number these companies publish is measured above the line where their largest expense sits.
The industry that already ran this play
One industry has already run this argument to its conclusion, in public, with court records, and it is the same industry currently rationing chips to the AI labs.
In the mid-1990s more than a dozen companies made DRAM at volume. By 2005 there were nine with meaningful share. Qimonda filed for insolvency in January 2009. Elpida filed in February 2012 and Micron bought it. Today three companies hold about 87 percent of the market.
Through the downturns, the survivors kept building capacity while their competitors cut, and prices stayed on the floor long enough to finish the weak ones off. Then the market consolidated, and the three survivors were caught running a criminal price-fixing cartel. Samsung paid a $300 million fine, Hynix $185 million, Infineon $160 million. Executives went to prison. The European Commission fined ten producers €331 million in 2010 for the same conspiracy. Micron got immunity for reporting it.
Now look at where that industry sits in 2026. In the first quarter of 2023, SK hynix posted an operating loss of 3.4 trillion won, its worst since 2012. In fiscal Q3 2026, Micron reported record results: $41.46 billion in revenue at an 84.9 percent non-GAAP gross margin. Conventional DRAM contract prices rose 93 to 98 percent quarter over quarter in the first quarter of 2026.
Seventeen years separate Qimonda’s insolvency filing from Micron’s 84.9 percent gross margin. The three companies that came through that stretch are the same three now rationing high-bandwidth memory to Nvidia, which is what part three of this series was about. The industry setting the ceiling on the frontier already ran the play the frontier is suspected of running.
Two details keep me from carrying that analogy further than it goes. China’s CXMT went from 4 percent of DRAM share to 10 percent in a year and IPO’d in July 2026 with the stock up 466 percent on debut, so high prices are already pulling in the entry that erodes them. And LCD panels ran the identical script, cartel convictions included, and then Chinese manufacturers flooded the market until Samsung and LG exited the business entirely.
Pricing power after consolidation is real. It is not permanent, and nobody has ever held it by being asked nicely.

Does moving on-prem fix it
The obvious answer to a rising bill is to stop renting. The arithmetic is harder than it was a year ago, and the reason is the subject of part three.
An Nvidia RTX Pro 6000 Blackwell opened pre-orders around $8,565 in early 2025. Nvidia’s own marketplace listed it at $13,250 in June 2026 and $16,000 in August. Commentary quoted in the Tom’s Hardware report attributed roughly 8 percent of the increase to actual cost and the rest to demand. Apple pulled the 512GB Mac Studio configuration in March 2026 over memory supply and will not bring it back until late October, above $10,000.
The break-even, in the most careful practitioner analysis I found, sits near 2 million tokens a day. Below one million, the API wins. Above ten million, owned hardware pays back in six to twelve months. Everything between is a judgment call about utilization, and utilization is where these business cases die. Hardware idle twenty-two hours a day costs more per token than any API rate on the market.
The line worth remembering from that analysis is that the MLOps hire costs more than the GPU. A $160,000 engineer against a $5,000 to $15,000 card is not a close comparison, and most spreadsheets leave the engineer out.
What has genuinely improved is the model side. OpenAI’s gpt-oss-120b carries 117 billion parameters with 5.1 billion active and fits on a single 80GB GPU. Epoch measures the best open-weight models trailing the best closed ones by about four months, or 8 points on its capability index. Open weights are no longer a data center problem.
The strongest argument for owning your own model has nothing to do with price. OpenAI’s deprecation policy gives at least six months’ notice for generally available models and as little as two weeks for preview models, and a large retirement wave lands on October 23, 2026. If you have production work sitting on an endpoint someone else can retire, you have an availability problem that no amount of cost optimization addresses.
What I would do about the bill before 2027
Four things, in the order I would do them.
Measure cost per completed task, not cost per million tokens. Every vendor quotes the second number and nobody is billed in it. A model that costs three times more per token and finishes the job in a quarter of the calls is cheaper, and you cannot see that on a rate card.
Turn on prompt caching. Anthropic charges 0.1 times the base input rate for a cache hit, and less on newer models. Writes cost 1.25 or 2 times depending on how long you hold the cache. For anything with a stable system prompt or a repeated document, this is the largest single lever available and it requires no negotiation.
Use the batch APIs where latency does not matter. OpenAI and Google both discount 50 percent for a 24-hour completion window. Overnight classification, summarization and enrichment jobs have no business running at interactive rates.
Route by task rather than by habit. Ramp’s panel already shows companies doing this, and it is the reason frontier share of tokens fell eight points in a month.
Between caching at a tenth of the input rate and batch at half, a workload with stable context and tolerant latency can absorb a doubling of list prices without anyone reopening a contract.
Where this lands
The strict version of the race-to-the-bottom argument does not survive the margin data, and I am not going to make it. The version that does survive is narrower and, for anyone holding a budget, worse. Three companies now take 88 cents of every enterprise API dollar, up from 69 cents in 2023. The open-weight share that was supposed to hold the price floor fell from 19 percent to 11 over the same two years. And OpenAI’s chief executive spent July publicly offering to sell at a quarter of a competitor’s price, which is what a company buying position sounds like.
Meanwhile the price of the thing you actually buy, a finished task, has been climbing the whole time. The increases that have already landed came through as quota resets, expired promotions, surcharged context windows, priority tiers and reasoning tokens, rather than as a number on a page.
Sanjay Mehrotra told Micron’s investors in June that multi-year customer agreements “will significantly enhance the durability and predictability of Micron’s strong financial performance.” He runs one of three companies left standing in a business that once had a dozen, at an 84.9 percent gross margin, three years after his closest competitor lost 3.4 trillion won in a single quarter. Durability and predictability is one way to say it.