Claude Sonnet 5’s Price Rise Comes Twice
A 50% jump on 1 September, and a tokenizer that already charges you more for the same words.
If you build on Claude Sonnet 5, your bill goes up on 1 September. The headline rate rises 50%, from $2/$10 to $3/$15 per million input and output tokens, as what Anthropic calls “introductory pricing” expires. This is the disclosed part. It sits in a note on Anthropic’s own pricing page, and it is the smaller of the two increases you are actually paying.
The larger one is quieter and already live. Sonnet 5 ships with a newer tokenizer that, by Anthropic’s own admission, cuts the same text into roughly 30% more tokens than the one Sonnet 4.6 used. The price per token can stay flat and your bill still climbs, because the identical prompt now buys you more tokens to pay for.
Neither move is hidden, exactly. Both are written down. But between an “introductory” rate that was always going to rise and a tokenizer that inflates the very unit you are charged by, the number on the pricing page has become one of the least reliable guides to what Sonnet 5 will cost you. This is the current state of the art in AI pricing: technically transparent, practically opaque.
The rise Anthropic tells you about
Start with the honest, above-the-line increase. Anthropic’s pricing documentation lists Sonnet 5 at $2 per million input tokens and $10 per million output tokens “through August 31, 2026,” then $3 and $15 “starting September 1, 2026.” A footnote spells it out: “Introductory pricing of $2/$10 per million input/output tokens is in effect through August 31, 2026, after which the standard pricing of $3/$15 per million input/output tokens will take effect.” The Batch API tier moves in lockstep, from $1/$5 to $1.50/$7.50. The same deadline applies whether you call Claude first-party or through Amazon Bedrock, Google Vertex AI or Microsoft Foundry.
To be fair, this is a normal, disclosed launch discount. Introductory pricing is not a dark pattern; it is a promotion with an end date, and you got months of a real 33%-off rate for it. Plenty of software is sold this way, and a company is entitled to charge standard price for a standard product. If this were the whole story, it would be a footnote in the literal sense: a scheduled reversion, announced in advance, that a competent finance team plans around.
The rise that rides in the tokenizer
It is not the whole story. A large language model does not read characters; it reads tokens, the sub-word chunks it slices your text into, and it bills you by the token in both directions. Change how the text is chopped and you change how many tokens the same passage becomes — and therefore the price, even if the price per token never moves.
Anthropic says plainly that this is what happened. Its pricing note reads: “Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.” Sonnet 5 is on the new tokenizer. Sonnet 4.6 is on the old one. The technology outlet The Decoder, which reported the pattern, cited community analysis measuring a 37.4% jump in tokens per request — higher than the “approximately 30%” on the label.
Give the good-faith reading its due: a tokenizer that produces more, smaller tokens can genuinely help a model reason and follow instructions, and Anthropic frames the change as a performance improvement rather than a billing one. That may well be true. It is also true that the improvement arrives attached to a higher token count on every request you pay for, and the two are not separable at the invoice.
There is one asymmetry worth flagging between the two increases. The introductory rate was time-limited, but at least it was a number you could see and diarise; when it ends, it ends cleanly. The tokenizer is different in kind. It is a fixed property of the model, not a promotion, so there is no version of Sonnet 5 that bills you on the old, leaner token count. You cannot wait it out, opt out, or negotiate it — if you want the model, you buy its tokenizer, and its tokenizer costs you more per sentence than the last one did.
Why “unchanged” is not unchanged
Here is where the sticker misleads. After 1 September, Sonnet 5 and the older Sonnet 4.6 both read $3/$15. Identical, apparently. But Sonnet 5 turns the same prompt into roughly a third more tokens, so for the same text you pay roughly a third more than 4.6 would have charged — at a rate the page insists has not changed. The increase is real; it is just relocated from the price column to the quantity column, where pricing pages do not look.
The effect compounds because Sonnet 5 is more agentic, spinning through more tool calls and more reasoning per task, which burns still more tokens before you reach the tokenizer question at all. The Decoder found an average task on Sonnet 5 costing $2.29 against $1.97 on the pricier-on-paper Opus 4.8 — a model whose per-token rate is higher yet whose real-world task came out cheaper. When a “$3” model can cost more per job than a “$5” one, the sticker has stopped doing the job stickers are for.
For anyone on a fixed Claude subscription rather than the metered API, the same dynamic arrives wearing a different hat: more tokens per task means fewer tasks per window before you hit a cap, which is its own quiet price rise. We have described how those limits are already hard enough to reason about; a tokenizer that enlarges every request does not make the arithmetic friendlier.
The pattern, not the incident
This is worth stating carefully, because intent is hard to prove and we do not assert it. What we can describe is the effect and the incentive. By The Decoder’s account, the same manoeuvre accompanied Opus 4.7, where the newer tokenizer was measured inflating token counts by between 1.325 and 1.47 times. Do it once and it is a tokenizer upgrade. Do it across a model line while the headline rates sit still, and it reads as a method: keep the advertised price legible and reassuring, and let the real increase accrue in the unit nobody audits.
We have written before about how token pricing looks transparent and predicts your bill badly, and how a flat subscription can quietly turn into a meter. This is the same species of problem seen from the API side. The per-token price is a genuine, published number. It is also, increasingly, the wrong number to plan with, because the count it multiplies is a moving target the vendor controls.
What it means for your bill
None of this requires outrage. It requires measurement, because the published figures will not do the estimating for you. If you run anything non-trivial on Sonnet 5, a short list before 1 September earns its keep:
- Re-measure on your own workloads. Run a representative batch through Sonnet 5 and Sonnet 4.6 and compare actual token counts, not advertised rates. The gap on your text is the number that matters, and it will not be exactly 30%.
- Treat the sticker as a floor, not an estimate. Forecast the 50% reversion and the tokenizer inflation together; they stack, and they land at the same time as the standard rate.
- Lean on the discounts that are real. Prompt caching and the Batch API remain the honest levers for cutting a Claude bill; use them before you assume the model is simply too dear.
- Benchmark the alternatives on cost-per-task, not cost-per-token. As Opus 4.8 showed, a higher rate can mean a cheaper job. Price the work, not the unit.
- Watch the deadline, not the announcement. The change is scheduled, so the surprise is avoidable — but only if you diarised 1 September rather than trusting that “same rate” meant same cost.
The price of transparency theatre
To Anthropic’s credit, all of this is documented. The reversion date is on the page; the tokenizer change is in a note; a diligent reader can reconstruct the true cost. That is genuinely better than the alternative, and better than several rivals manage. We would rather have the footnote than not.
But disclosure and legibility are not the same thing, and the distance between them is exactly where the money moves. A price you can only understand by reading two notes, doing arithmetic across tokenizers and measuring your own workload is transparent in the way a mortgage contract is transparent. The fix on the reader’s side is unglamorous and effective: keep receipts, measure your real usage, and stop letting a reassuring rate on a pricing page stand in for the bill it no longer predicts. The number went up twice. Only one of the increases wanted to be seen.
Frequently asked questions
When does Claude Sonnet 5’s price go up?
On 1 September 2026. The introductory rate of $2 per million input tokens and $10 per million output tokens reverts to the standard $3/$15, a 50% increase on both. Anthropic states this in a note on its pricing page.
Is the per-token price the whole story?
No. Sonnet 5 uses a newer tokenizer that, by Anthropic’s own account, produces about 30% more tokens for the same text than Sonnet 4.6’s. Since you pay per token, the identical prompt costs more even when the rate is unchanged.
Does this affect Claude on AWS Bedrock, Google Vertex or Microsoft Foundry?
The introductory pricing and the 31 August deadline apply across Anthropic’s first-party API and the partner cloud platforms, per its pricing documentation. The tokenizer is a property of the model, so it travels everywhere Sonnet 5 does.
Is Sonnet 5 actually more expensive than Sonnet 4.6?
At list, after 1 September both show $3/$15. But Sonnet 5’s tokenizer emits more tokens per request and it behaves more agentically, so for the same task it tends to cost more in practice despite the identical headline rate.
What should I do before 1 September?
Measure your real token usage on Sonnet 5 versus 4.6 on your own prompts, fold the tokenizer change into your forecasts rather than trusting the sticker rate, and lean on prompt caching and the Batch API where you can.