Every few weeks I get the same call. A company has moved an AI pilot into production and the bill is two or three times what they expected. The vendor's pricing page made it look cheap, and the proof of concept ran on a few hundred dollars a month. Now they're looking at five figures and trying to work out where it went.

The odd part is that per-token prices have fallen over the last 18 months. The spike doesn't come from the headline numbers. It comes from architecture choices that never show up on the pricing page.

Where the money goes

When I audit an AI workload, the bill usually splits into six buckets. Most teams underestimate at least three of them.

1. Context, the silent multiplier

Input tokens dominate, and that surprises almost everyone. A typical enterprise agent sends a system prompt, a knowledge base excerpt, the chat history and sometimes a set of tool definitions. That's often 8,000 to 20,000 tokens of input for an answer of maybe 200 tokens.

At $3 per million input tokens and $15 per million output, the input alone can be 80% of each call's cost, and it's sent again on every turn of the conversation. Caching helps a lot, but most teams haven't switched it on.

2. Agent loops

An agent that uses tools doesn't make one model call per request. It often makes five, ten or twenty, and each one sends the full context plus the latest tool result back to the model. One agent interaction can burn through 100,000 tokens before anyone notices.

This is the bucket I see underestimated most. The maths people do at design time assumes one call per request, and production is rarely that tidy.

3. Retrieval infrastructure

Vector databases, embedding APIs, and the storage and re-indexing needed to keep a knowledge base fresh. It's small per query, but it adds up at scale and it's a standing infrastructure cost rather than a per-call one.

4. Observability and evaluation

Anything in production needs logging, tracing and ongoing evaluation. Tools like LangSmith and Helicone, or whatever you build yourself, cost money. The eval runs you do on every model change cost real tokens too, sometimes more than production traffic while development is busy.

5. More than one model

Most serious deployments now use several models. A cheap one handles routing or first-pass triage and a flagship takes the hard cases. Some teams add a third for verification. Each adds latency and cost, though it usually still beats sending everything to the flagship.

6. People

Cost models often leave out the people who curate prompts, review outputs, tune evaluations and handle escalations. For any AI system that staff or customers rely on, that's a real operating cost. It's labour rather than cloud, and it decides whether the system improves or drifts.

What brings the cost down

Prompt caching

Anthropic, OpenAI and Google all support it now, so turn it on. Savings on repeated system prompts and knowledge base context are typically 60–90%. For most teams in 2026 it's the easiest five-figure saving on the table, and most haven't taken it.

Model routing

Don't send every query to your most expensive model. Build a simple classifier (often a small language model) that decides whether a query needs the flagship or whether a smaller model will do. On a procurement assistant I worked on recently, routing cut total inference cost by about 70% with no quality drop we could measure.

Output budgets

Set max_tokens aggressively. If your use case never needs more than 500 tokens of output, a cap stops a runaway generation from costing more than it's worth. Everyone agrees with this and hardly anyone does it.

Batch where you can

For offline work like summarising, classifying and enriching data, most providers offer batch APIs at half price or less. If it doesn't need an answer in real time, send it through the batch path.

Measure tokens like you measure latency

Every production AI system should track token usage as a first-class metric, broken down by feature, user segment and model. Without that, cost work is guesswork. With it, the expensive paths are obvious within days.

What the well-run organisations have in common

Companies with AI costs under control treat AI spend the way they treat cloud spend, with FinOps practices, monthly reviews and ownership at team level. They've usually built a thin internal layer over their model providers, so they can route, cache and switch without rewriting application code. And someone's job includes watching the bill and asking why it moved.

The ones without control treat cost as the vendor's problem and wait for the next price cut. The price cut usually does arrive. It rarely fixes the architecture underneath.

Where to start

Start with visibility. Tag every model call with the feature that triggered it, the user segment, and tokens in and out. A week of that data usually points to one or two paths eating most of the budget. Fix those first and leave the rest for later.

If you'd like someone to go through where your AI spend is really going, I'm happy to take a look.