Tokenised pricing: why your AI budget will be wrong
Consumption billing in a unit almost nobody reasons about intuitively. Where the overruns come from, the four discounts most buyers never ask for, and why term length matters more than rate.
General information only. Verify before you act. This article is published for general information and is not advice. It does not take account of your circumstances, your agreements or your obligations, and no advisory or client relationship arises from reading it.
Third party product names, licensing structures and prices change frequently and without notice. Any figure, range or percentage here is indicative, may be out of date at the time you read it, and is not a prediction of any result in your organisation. Carry out your own due diligence before acting. Verify everything against the supplier's current published documentation and your own contracts, and take advice from a suitably qualified professional where the decision warrants it.
Valence Dynamics accepts no liability for any loss arising from reliance on this article. See our Terms of Use.
A note on figures, current as of 28 August 2026. Deliberately, this article contains almost no headline model prices. Per-token rates on the major providers changed repeatedly through 2026, in both directions, and any specific number would be stale within weeks. What does not change nearly as fast is the structure of the billing, which is where the procurement decisions actually sit. Get current rates from the provider's live pricing page on the day you model.
Tokenised pricing is the commercial model behind most AI products you are being sold. It is not a new idea, it is consumption billing with an unfamiliar unit, but the unfamiliarity is the problem. Most buyers can reason about a per-seat price. Very few can reason about a per-token price, and the gap between those two abilities is where the budget overruns live.
What a token actually is, commercially
A token is a fragment of text, roughly three quarters of a word in English, though it varies by language and content type. Code, structured data and non-Latin scripts consume more tokens per unit of meaning than plain English prose does.
What matters commercially is that you are billed on both directions of the exchange. Input tokens cover everything sent to the model, and output tokens cover everything it returns. These are almost always priced differently, with output typically costing several times more than input.
The consequence people miss: input is not just the question your user typed. It includes the system instructions, any retrieved documents, and in a conversation, the entire history resent on every turn. A ten-turn conversation can bill the early turns ten times over.
Why the estimate you were given is usually wrong
Vendor and internal estimates tend to fail in the same predictable ways.
Conversation history compounds
Because context is resent with each turn, cost per conversation grows faster than linearly with conversation length. A pilot benchmarked on single-question interactions will materially understate production cost, where users hold long sessions.
Retrieval inflates input
Any system that pulls in documents to ground its answers is injecting large volumes of input tokens the user never sees. The retrieved context frequently dwarfs the actual question. This is the single largest source of variance between pilot and production numbers.
Reasoning output is invisible but billable
Models that perform extended internal reasoning generate tokens that never appear in the answer. These are typically billed as output, at the higher rate. A short visible response can sit on top of a great deal of billed generation.
Agents multiply everything
An agentic workflow makes multiple model calls per user request, each carrying its own context. One user action can become a dozen billed exchanges. If you are being sold agents, per-request cost is the wrong unit and per-task cost is the right one.
Failure and retry are billed
Failed calls, retries, timeouts and outputs the user rejects and regenerates all consume tokens. Budget models built on successful-path assumptions understate reality.
The four discounts most buyers do not ask for
Before negotiating rate, exhaust the mechanisms that are already published and rarely volunteered. These are structural features of how the major providers bill, and they are available without a commitment conversation.
Cached input
Repeated context, typically a long system prompt or a stable document set, can be cached and re-read at a substantial discount to the standard input rate. On the major providers this is roughly a tenth of the standard rate, though the multiplier varies. There is a wrinkle worth checking: cache writes may be billed separately, and on some platforms at a premium to the standard input rate. Microsoft indicated cache write billing on Azure would begin from around 21 August 2026, having previously not been charged, so a caching architecture that penciled earlier in the year may need re-modelling.
Batch processing
Work that does not need an immediate answer can be submitted asynchronously, typically at around half the standard rate. Reporting, classification, enrichment, summarisation of overnight data are all natural candidates. Most organisations run everything synchronously by default because nobody asked which workloads actually need to be.
Model tier selection
The spread between the cheapest and most expensive tier from a single provider currently runs to two orders of magnitude or more. Classification, extraction and routing tasks generally run fine on the cheapest tier. The disciplined pattern is to start on the cheapest model that could plausibly work and escalate only where evaluation demonstrates a quality gap, rather than defaulting everything to the flagship because that is what was demonstrated.
Long context is a separate rate
Past a defined threshold, per-token rates on long inputs can roughly double. If your architecture routinely sends very large contexts, you may not be paying the rate on the pricing page you read.
These stack. Caching a large stable system prompt, batching the non-urgent workloads, and moving classification tasks down to a cheap tier are three independent changes to the same bill. Applied together they routinely cut spend by a substantial multiple, with no change to the product or the output quality.
Two platform issues specific to enterprise deployments
The same model can appear on two invoices
Where an organisation uses both a hyperscaler-hosted deployment and the provider's direct API, the identical model can appear on the cloud bill and on the provider's bill in the same month, for the same product feature, at different rates. Deployment type matters too: on Azure, globally distributed deployments generally track the provider's list rates while data-residency-constrained deployments cost more. If you have a residency requirement, you are paying for it, and that should be a deliberate decision rather than a default.
Retrieval carries its own meter
A retrieval-augmented feature typically bills on at least two lines: the model tokens, and the search or vector service that supplies the context. On Azure, AI Search is a separate meter entirely. Any cost model that counts only tokens for a RAG application is understating the position before the first question is answered.
The structural problem with committing
Vendors will offer a discount for committing to a volume of consumption up front. The discount is real. The difficulty is that you are being asked to forecast a quantity you have no historical basis for estimating, in a unit you do not natively think in, for a product whose usage pattern will change as your users learn what it can do.
This is the same structural trap as an oversized cloud commitment, with one aggravating factor: at least with compute you had years of utilisation data. With AI consumption you frequently have a three-week pilot.
There is a second dynamic, and 2026 has demonstrated it in both directions. Rates on a given model family have moved sharply and unpredictably. One major provider roughly doubled per-token pricing on a flagship line in April 2026, then cut prices on newer tiers by 20 and 80 percent respectively at the end of July. Over the same period the flagship tier across competing providers converged to broadly similar rates.
The lesson is not that prices always fall. It is that they move faster than your contract term, in ways nobody forecasts reliably. That makes term length more important than discount percentage, and makes a price review or benchmarking clause more valuable than an extra few points off the rate.
Questions to ask before signing
- 01What is the cost per completed business task, not per call?Insist the vendor model a realistic end-to-end task including retrieval, retries and multi-step calls. If they can only quote per million tokens, they have not done the work.
- 02What happens if we exceed the commitment?Overage rates are sometimes at list price, wiping out the discount. Establish the rate and whether there is a burst allowance.
- 03What happens if we do not reach it?Is unused commitment forfeited at period end, or does it roll forward? Forfeiture converts your discount into a penalty for over-forecasting.
- 04Can we move down as well as up?Most agreements make growth easy and contraction impossible. Negotiate a reduction right, even a limited one, before signature. It is very hard to get afterwards.
- 05What is the price protection if published rates fall?Rates moved twice on major providers during 2026, including an 80 percent cut on one tier. Ask for published price reductions to flow through, or at minimum a mid-term benchmarking review. This is worth more than a few points off the headline rate.
- 06What granularity of usage reporting do we get?You cannot govern what you cannot see. Insist on consumption broken down by application, team and use case, not one aggregate line. Without this, chargeback and optimisation are impossible.
- 07Are we billed on the model we need, or the one we were demoed?Capability tiers differ in price by an order of magnitude. Many workloads run fine on a cheaper tier. Establish which tier each use case actually requires.
- 08What are the caching and batching terms?Cached input runs at roughly a tenth of standard rates and batch at around half. Confirm the cache write multiplier separately, since it changed on Azure in August 2026.
What good governance looks like
Treat AI consumption exactly as you would treat cloud spend, because it behaves the same way. That means tagging usage by team and use case from day one, setting hard budget alerts rather than soft ones, running the cheapest model tier that meets the quality bar for each workload, and reviewing consumption monthly against forecast rather than discovering the variance at year end.
The organisations that control this cost are not the ones that negotiated the best headline rate. They are the ones that instrumented usage before it scaled.
Want this checked against your own estate?
A discovery call is free and takes 30 minutes. Tell us what you are spending and where you think the gaps are, and you will get an honest view of whether there is anything worth pursuing.
Get in touch