TL;DR
- A "token" is not a fixed amount of text. Each vendor's tokenizer cuts the same file into a different number of pieces, and you pay per piece. So $/Mtok is not comparable across vendors.
- Anthropic's newest tokenizer (Sonnet 5, Opus 4.8, Opus 5, Fable 5, Fable 5.1) produces ~30% more tokens from the same code than their previous one. The list price did not change. Opus 5 and Fable 5.1 were verified against real bills on 2026-09-02: same tokenizer, to the token.
- On identical files, it produces 1.36-1.73x GPT's token count. TypeScript is the worst case at 1.73x.
- In effective terms, Opus 4.8's and Opus 5's $5 / $25 behave like $7.50 / $37.50, and Fable 5.1's $10 / $50 like $15 / $75. Sonnet 5's $2 / $10 was announced as an intro price ending August 31; on September 2 it is still the live rate.
- This measures input tokenization only. Output is a bigger lever: on one identical task, Fable 5.1 wrote 2.3x the output tokens of Fable 5 and Sonnet 5 wrote 8.8x Gemini 3.7 Flash. Details near the end.
We counted the same bytes under every frontier tokenizer, using each vendor's own counting endpoints, and cross-checked the counts against real paid requests. Below are the numbers and what they do to the prices on the rate cards.

Why $/Mtok is not a comparable price
A model's bill is two numbers multiplied together:
cost = (tokens your content becomes) x (price per token)
The pricing page shows the second number and treats the first as a constant. It is not a constant. It depends on the model's tokenizer, and tokenizers differ a lot between vendors. Two models can list the same "$5.00 / 1M input tokens" and produce meaningfully different bills for the same paragraph, because one of them turns that paragraph into more tokens. Since nobody publishes tokens-per-content numbers, we measured them.
How we measured
We took 16 real fixtures: English prose, an HTML page, JavaScript, Python, TypeScript and Rust files, JSON tool schemas and tool results, Chinese chat and prose, symbol-heavy text, and our own agent system prompt. Each fixture was counted, byte for byte, with every model's production tokenizer:
- Anthropic models were counted with the official
count_tokensendpoint, which returns the same count Anthropic bills against. - OpenAI models were counted with the documented
o200k_basetokenizer viatiktoken. For the newest models we double-checked this against production: we sent real API calls to GPT-5.1, GPT-5.5, and GPT-5.6 Sol and compared the liveusagenumbers with the local count, using a long-minus-short delta to cancel the request framing. All three matchedo200k_baseexactly. - Gemini and Grok were counted with their providers' token-count endpoints.
GPT's o200k serves as the 1.00x reference throughout, mainly because it has been frozen and publicly documented for over two years, while Claude's tokenizer is the one that changed. DeepSeek and GLM are left out of the tables entirely: we only have rough characters-divided-by-four estimates for them, not real tokenizer counts, and this post is about measured numbers.
Finding 1: same list price, ~30% more tokens
Claude Opus 4.6 and Opus 4.8 have the same $5.00 / $25.00 list price. What changed between them is the tokenizer. Sonnet 4.6 and Opus 4.6 use the old one; Sonnet 5, Opus 4.8, Fable 5 and, since this update, Opus 5 and Fable 5.1 use the new one. The table counts the same bytes with both, on Anthropic's own endpoint:
| Content | Old tokenizer | New tokenizer | Change |
|---|---|---|---|
| English prose (2,115 chars) | 476 | 636 | +34% |
| HTML page (3,195 chars) | 1,131 | 1,302 | +15% |
| JavaScript (1,933 chars) | 659 | 794 | +20% |
| Python (2,251 chars) | 831 | 1,022 | +23% |
| TypeScript (2,888 chars) | 898 | 1,178 | +31% |
| Rust (2,924 chars) | 1,019 | 1,312 | +29% |
| JSON tool schema (9,948 chars) | 2,631 | 3,306 | +26% |
| Our agent system prompt (42,661 chars) | 10,761 | 14,953 | +39% |
| Chinese prose (379 chars) | 435 | 433 | ~0% |
Weight those rows the way a real agent request is composed, which is mostly English system prompt, tool schemas, code, and JSON, and the new tokenizer comes out around +32% per request. The Chinese row barely moved, so the inflation is concentrated in English and code.
Sonnet 5's launch price, recalculated
Sonnet 5 launched at $2.00 / $10.00, down from Sonnet 4.6's $3.00 / $15.00, which looked like a price cut. Anthropic announced it as an intro price ending August 31, 2026. While it holds, the lower rate slightly more than covers the extra tokens, so Sonnet 5 works out a little cheaper than 4.6 for the same code. Update 2026-09-02: the rate card still shows $2.00 / $10.00 two days after the announced end, with no end date on the page. If it does return to $3.00 / $15.00, the extra tokens remain and the same work will cost about a third more than it did on Sonnet 4.6 at the same list price; the table below shows both.
Checking the counter against real bills
count_tokens is a prediction, so we also sent real paid requests with max_tokens: 1 and read usage.input_tokens, which is what invoices are based on. For the same content, Opus 4.6 billed 2,541 input tokens and Opus 4.8 billed 3,191, each matching its predicted count exactly. We ran the same check on Fable 5, the most expensive model in the lineup, and it billed 3,191 as well, identical to Opus 4.8. So Fable uses the same new tokenizer and there is no extra per-token markup hidden behind its higher list price. The whole verification cost about $0.08.
Re-verified 2026-09-02 for the two models Anthropic shipped since: Opus 5 billed 3,191 for the same delta, and Fable 5.1 billed 3,191 too. Fable 5.1's raw counts sit exactly 2 tokens above Fable 5's on every one of the 16 fixtures (117 against 115 on the short one, 3,308 against 3,306 on the long one), and a constant offset that cancels in the delta is request framing, not a tokenizer. So the rumour that 5.1 "eats more" is not about input: the same bytes cost the same tokens. Where it eats more is output, below. This pass cost $0.12.
Finding 2: the gap is widest on code
The cross-vendor table uses GPT's o200k as the 1.00x reference. Every cell is that model's token count for the identical file divided by GPT's, so 1.20x means 20% more tokens than GPT. Claude's new and old tokenizers are shown side by side:
| Content | Claude (new) | Claude (old) | Gemini 3.x Flash | Grok 4.5 |
|---|---|---|---|---|
| TypeScript | 1.73x | 1.32x | 1.16x | 1.05x |
| Rust | 1.58x | 1.22x | 1.19x | 1.05x |
| JavaScript | 1.52x | 1.26x | 1.23x | 1.11x |
| Python | 1.50x | 1.22x | 1.20x | 1.09x |
| HTML page | 1.36x | 1.18x | 1.08x | 1.04x |
| English prose | 1.40x | 1.05x | 1.01x | 1.00x |
| Chinese prose | 1.44x | 1.45x | 0.85x | 0.86x |
| Chinese chat | 1.53x | 1.55x | 0.91x | 0.92x |
The code rows sit well above the prose rows: TypeScript at 1.73x, Rust at 1.58x, JavaScript at 1.52x, Python at 1.50x, against 1.40x for English prose. Code is most of what a coding agent processes, so for that workload the 1.50-1.73x band is the relevant one.
Why is TypeScript the worst case? Because o200k is unusually efficient on it: about 4.24 characters per token, which looks like the result of training on a lot of web JavaScript and TypeScript, where camelCase identifiers and JSX patterns compress into single tokens. On Rust its efficiency drops to about 3.51 characters per token. Claude's tokenizer is roughly equally dense on both languages, so the gap is widest exactly where GPT is strongest.
Chinese behaves differently. Claude sits around 1.45-1.55x above GPT with both the old and the new tokenizer (435 vs 433 tokens against GPT's 300 on the prose fixture), so this is a long-standing property of the Claude family on CJK text, not something the new tokenizer introduced. Gemini is actually more efficient than GPT here, at 256 tokens. Which tokenizer costs you more depends on what you write.
The Gemini column now reads "3.x" because Gemini 3 Flash, 3.6 Flash and 3.7 Flash return identical counts on every fixture (789 / 459 / 2,541 on the TypeScript, prose and tool-schema files, counted 2026-09-02 on Vertex). One tokenizer across the generation, so a Gemini price change is a real price change.
What that does to the price
Multiply the list price by the measured divergence and you get an effective price for processing the same work. Divergence here is the blended multiplier for a typical English coding request, normalized to GPT's o200k:
| Model | List price in / out ($/Mtok) | Divergence | Effective in / out ($/Mtok) |
|---|---|---|---|
| GPT-5.1 | $1.25 / $10.00 | 1.00x (reference) | $1.25 / $10.00 |
| GPT-5.5 | $5.00 / $30.00 | 1.00x | $5.00 / $30.00 |
| GPT-6 Astra | $10.00 / $50.00 | 1.00x (verified 2026-09-05) | $10.00 / $50.00 |
| GPT-5.6 Sol | $4.00 / $20.00 | 1.00x (verified) | $4.00 / $20.00 |
| GPT-5.6 Terra | $2.00 / $12.00 | 1.00x (verified 2026-09-02) | $2.00 / $12.00 |
| Grok 4.5 | $2.00 / $6.00 | 1.03x | $2.06 / $6.18 |
| Gemini 3 Flash | $0.50 / $3.00 | 1.09x | $0.55 / $3.27 |
| Gemini 3.7 Flash (through 2026-12-31) | $0.75 / $3.75 | 1.09x (same tokenizer) | $0.82 / $4.09 |
| Claude Sonnet 4.6 | $3.00 / $15.00 | 1.14x (old tokenizer) | $3.42 / $17.10 |
| Claude Sonnet 5 (live rate, 2026-09-02) | $2.00 / $10.00 | 1.50x (new tokenizer) | $3.00 / $15.00 |
| Claude Sonnet 5 (if the intro ends) | $3.00 / $15.00 | 1.50x | $4.50 / $22.50 |
| Claude Opus 4.6 | $5.00 / $25.00 | 1.14x (old tokenizer) | $5.70 / $28.50 |
| Claude Opus 4.8 | $5.00 / $25.00 | 1.50x (new tokenizer) | $7.50 / $37.50 |
| Claude Opus 5 | $5.00 / $25.00 | 1.50x (verified 2026-09-02) | $7.50 / $37.50 |
| Claude Fable 5 | $10.00 / $50.00 | 1.50x (new tokenizer) | $15.00 / $75.00 |
| Claude Fable 5.1 | $10.00 / $50.00 | 1.50x (verified 2026-09-02) | $15.00 / $75.00 |
A few rows are worth a second look. Opus 4.6 and 4.8 share a list price but differ by about 32% in effective price, and Opus 5 inherits 4.8's. GPT-5.5, GPT-5.6 Sol and GPT-5.6 Terra share the tokenizer, so their list prices really are comparable with each other. Gemini Flash runs a slightly heavier tokenizer than GPT and still remains the cheapest option by a wide margin. Fable 5.1 has one line the table cannot show: its cache reads are priced at 0.025x base ($0.25 / Mtok) instead of the usual 0.1x, so on a long agent session, where cache reads are most of the bill, its effective per-token price falls well below Fable 5's despite the identical sticker.
For an independent data point: Ploy published a production migration to GPT-5.6 Sol this week and reported 1.70M input tokens against Claude Opus 4.8's 2.60M for the same builds, about 35% fewer. That is a whole-task bill rather than a tokenizer probe, so it also folds in model verbosity, but it points the same way.
What the input ratio does not capture
Everything above measures one thing: how many input tokens identical bytes become. A full agent task adds more variables on top, and they are big ones. How many output and thinking tokens does the model spend to reach the same result? How much context does the harness load per step? How often does it call tools or spawn subagents? How does the provider price cache reads and writes?
Two consequences are worth spelling out. First, cache traffic is billed per token too, so a tokenizer that produces 32% more tokens also makes every cache write and every cache read about 32% more expensive, and on long agent sessions cache reads are most of the bill. Second, whole-task costs can diverge far more than 1.73x in either direction once verbosity and thinking are folded in. When people report that one model "uses 2-4x the tokens" of another on agent work, that can be true for their setup even though the pure input tokenization gap in our fixtures never exceeded 1.73x. The two numbers measure different layers.
The output layer, measured
We now have the second layer on one identical prompt. In the MacBook drawing benchmark, each model was asked for the same SVG at its top reasoning effort, with the output budget set to the model's own API ceiling. Output tokens spent on one drawing, thinking included:
| Model | Output tokens, one drawing | vs Fable 5 | Output ceiling |
|---|---|---|---|
| Claude Sonnet 5 | 121,713 | 2.9x | 128,000 |
| Claude Fable 5.1 | 99,072 | 2.3x | 128,000 |
| GPT-6 Astra | 45,836 | 1.1x | 128,000 |
| Claude Opus 5 | 74,061 | 1.7x | 128,000 |
| Claude Fable 5 | 42,560 | 1.0x | 128,000 |
| GPT-5.6 Terra | 29,834 | 0.7x | 128,000 |
| Gemini 3.7 Flash | 13,863 | 0.3x | 65,536 |
Same tokenizer, same sticker, and Fable 5.1 spends 2.3x the output tokens of Fable 5 to answer the same prompt at the same effort - because it thinks longer, and thinking is billed as output at $50 / Mtok. That is where "5.1 eats more" is true, and it is a bigger multiplier than anything in the input tables above. It also means the ceiling matters: 128,000 output tokens is the most any Claude or GPT-5.6 model will write in one answer (Gemini Flash stops at 65,536; only DeepSeek V4 Flash goes higher, at 393,216), and Fable 5.1 used 77% of it on one laptop. Our first pass at that benchmark capped output at 64K and read the cut-off as a model failure; it was ours.
By content type, the input-side range we measured for Claude's new tokenizer against GPT's o200k is: prose, HTML, and JSON at 1.36-1.42x; code at 1.50-1.73x with TypeScript on top; Chinese and symbol-heavy text at 1.44-1.53x. We put TypeScript in the title because it is both the top of the range and the thing a coding agent processes all day, not because the whole world is 1.73x.
The whole per-task layer lives in that follow-up: 16 models given one identical drawing task across the full effort ladder, 98 one-shot runs, every attempt priced from the providers' own usage numbers. The same drawing ranged from $0.004 to $4.95 depending on model and reasoning effort. That experiment is here: MacBook SVG Benchmark: 16 AI Models, Ranked by the Laptop They Draw.
How to compare model prices
- Compare on your own content. Your language and file types set the multiplier, so run a representative sample through each tokenizer before trusting a rate card.
- Treat a tokenizer change as a price change. When a vendor ships a new model at the same list price, check whether the tokenizer moved. Opus 4.6 to 4.8 is a ~32% increase with no line item on any invoice.
- Measure dollars per completed task, not dollars per token. That single number folds in tokenization, verbosity, thinking, and caching at once, and the provider's
usagefield gives you the ground truth to compute it. - $/Mtok is still useful as an opening line. It just is not sufficient, and it is not comparable across tokenizers. Vendors could fix this tomorrow by also publishing prices per byte; until someone does, the conversion work is on you.
None of this makes one model universally right. GPT-5.x is the token-lean choice on English and code, Gemini 3 Flash is remarkably cheap in effect, and Claude models earn their place on quality even when they cost more tokens to run. Just make sure the price you compare is the one you actually pay, after the tokenizer.
Sources: Anthropic pricing (anthropic.com/pricing), OpenAI pricing (platform.openai.com/docs/pricing), Google Gemini API pricing (ai.google.dev), xAI pricing (docs.x.ai). Token counts come from Anthropic's count_tokens endpoint, OpenAI's o200k_base (verified against live API usage), and the Google and xAI count endpoints. No text was generated to produce these counts.
The measurements are ours; the prose was drafted with AI assistance and edited by a human. Updated 2026-07-14 with a TL;DR, tighter wording, and a section on what input tokenization does not capture, based on reader feedback. Updated 2026-09-02: Opus 5, Fable 5.1 and GPT-5.6 Terra verified against real bills; Gemini 3.7 Flash counted; Sonnet 5's intro price re-checked against the live rate card; the output layer measured on one identical task.
Playcode keeps every one of these models one click apart, so you can run the same prompt on two of them and compare the result that matters, the app it builds, instead of arguing about a sticker. Try it at playcode.io.