← All Thought Starters
AI Tech Talk

The Price Per Token Went Down. Why Didn't Your Bill?

AI is billed in tokens, and a new model can split the same text into more of them. That's why a cheaper per-token rate doesn't always mean a cheaper document.

When Anthropic released Claude Sonnet 5 at the end of June, the per-token price was lower than the model it replaced. Plenty of teams switched and expected their monthly API costs to fall. For some of them, they didn't. The reason is a part of the system most people never think about, called the tokenizer.

What a tokenizer does

AI models don't read words the way you do. Before a model sees your text, a tokenizer chops it into pieces called tokens. A token might be a whole short word, part of a longer word, a space or a punctuation mark. Every AI provider bills by the token, for what you send in and for what the model writes back.

Different models can use different tokenizers, and each one splits text its own way. The same paragraph can come out as a different number of tokens depending on which model reads it.

What changed with Sonnet 5

Anthropic's own documentation for Sonnet 5 says its new tokenizer produces roughly 30% more tokens for the same input text than Claude Sonnet 4.6 did, and that the exact increase depends on the content. The same documentation says the per-token price is lower, but that the cost of an equivalent request does not drop in direct proportion.

Put simply, each token got cheaper, and each document now contains more tokens. Whether you save money depends on which of those effects is bigger for the kind of text you process. It also means a limit set in tokens, such as a maximum response length, now covers less text than it used to. A setting that worked on the old model can cut off answers on the new one.

How to check your own costs

Take a sample of the real documents or prompts your workflows handle, perhaps twenty of them. Run them through both models, or use the provider's token-counting tool, and compare the total cost of the same work. That number is the one that matters, and a pricing page can't tell you what it is. Do this before switching any automated process that runs at volume.

The caveat

Token counts are only part of the picture. A model that gets the answer right on the first try can cost less overall than a cheaper model you have to run twice, and staff time spent fixing output is a real cost too. Compare the cost of finished, correct work.

The takeaway

Treat the per-token price as a starting point. When you change models, measure what your own workload costs on the new one, and recheck any limits you set in tokens.

Thinking about how this applies to your business?

Book a conversation