Why your agent's token bill keeps growing, and how to audit it

An agent pipeline's token bill can double in two weeks with no new customers, no new workloads, and no model change. It looks like a pricing bug. It almost never is.

The numbers in this post come from a real production agent setup that went from about 110 million tokens a day to over 500 million on its peak day this month, and from the audit that followed. The setup's details matter less than the shape of the causes, because the same three causes show up in almost every agent pipeline once you measure them.

If you're running agents of your own, some version of these questions has probably crossed your mind:

  • Why is my agent's token bill so high all of a sudden?
  • Why does a single task cost more than it did last month?
  • Does a longer conversation really cost more, or is that my imagination?
  • How do I cut my agent's token costs without crippling it?
  • What's the cheapest way to give an agent long-term memory?
  • How do I tell which parts of my setup are worth pre-loading?

All six have the same root: one human message is not one billable event. Each message kicks off an agent loop: think, call a tool, read the result, think again. Every loop iteration is one LLM request that re-reads the entire conversation so far. In the pipeline we audited, 285 messages on the peak day became roughly 3,000 billable requests.

That means context size is a multiplier on everything you do. Here is the math, with the audited pipeline's October numbers preloaded. Drag the sliders to your own setup.

Context-cost simulator · per-request compounding
200
5
12
15
30
—
cumulative tokens re-read
—
estimated cost (cache $0.27 / input $3.00 per 1M)
—
context size at final iteration
Model: each iteration re-reads current history (cache rate) and appends new content (input rate). History grows linearly: H(n) = H0 + n·Δ. Assumes ~250 tokens per KB of mixed English/code; prices shown are public list prices for a popular open-weight model — your provider's rates apply the same shape.

The compounding is the whole game: every loop iteration makes history bigger, and bigger history makes every subsequent iteration more expensive. With that in hand, the three causes.

Cause 1: Volume (the honest one)

The least interesting cause usually does most of the work. In the audited pipeline, one active conversation a day became 17 to 22 parallel ones during a holiday stretch, and 95 LLM calls a day became roughly 3,000. Nothing was wrong. More work was simply being done.

Volume is easy to spot and easy to excuse, which is why it often hides the other two causes. Always separate it first.

Cause 2: Thinking effort, a same-day A/B

Many clients expose a thinking-effort setting. In the audited pipeline, work had quietly shifted toward the max setting over the same period. Because both settings were in use on the same days, the two could be compared directly. Median context replay per request:

Date high-effort sessions max-effort sessions
Oct 2 241k tokens 342k tokens
Oct 4 252k tokens 413k tokens

Same days, same kind of work, 40–65% more replay per request under max. Longer thinking blocks accumulate into conversation history, and history is what every subsequent request re-reads. The practical rule: reserve the strongest setting for genuinely hard problems — architecture, attribution, design review — and run execution and bookkeeping one notch down.

Cause 3: Mechanism density (the one that creeps)

The least obvious cause, and the one that never stops growing by itself.

Agent setups accumulate tooling: a registry to grep before building anything, a closeout checklist, a reconciliation scan, a delivery rubric. Each addition is small. But every added protocol step means more file reads, every file read lands in conversation history, and history is re-read by every later request. In the audited pipeline, the median fresh input per request jumped from 3.3k to 8.4k tokens the day a batch of new protocols went live, with zero change in what the agent was asked to do. Governance files were read 277 times in three days.

Per-request replay, the part of the context re-read from cache on each call, went from a ~112k median in late September to 299–386k in early October. None of these changes is a bug. Together they are a multiplier on every single request, and unlike volume, this one does not go away when you work less.

What actually moves the number

Four techniques, in the order they tend to pay off.

Score context by frequency × size × residency. File size tells you about storage; context cost tells you about reading. A 17KB rulebook loaded into every session costs more than a 1MB archive read once a month. Anything in the working set is re-read by every request. Score everything your agent pre-loads by how often it loads, how big it is, and whether it stays resident — that ranking is your trim order.

Gate every downgrade. Anything can leave the pre-load path and become read-on-demand, but only after one question is answered: which scenario will need this, and what catches that scenario? A protocol that only matters when handling people, for example, needs something that routes people-handling tasks to it. If a scenario has no catcher, the file stays pre-loaded. Skipping this question is how red-line files get quietly optimized away.

Tier the loading. "Load everything at session start" is two different needs wearing one coat: the rules that must always apply, and the state that's only needed when someone asks. In the audited setup, splitting a 48KB full read into a 14KB warm start plus an on-demand full read removed most of the per-session cost without removing any rule.

Trim the injected index. Most clients inject a catalog of available skills or tools into every system prompt. In the audited setup that catalog had reached 241 descriptions, 161.6KB, 80% of the system prompt, with one plugin family contributing 133 entries. Disabling five unused plugin packs, a flag flip with files intact, took the index to 29.1KB and the total system prompt from 200KB to 63KB, measured on a fresh session. One honest caveat on scale: an injected index rides the cache path, so at cache pricing this was worth 12–15% of per-request replay, not the two-thirds the byte count suggests. Measure savings in tokens at the price tier they actually ride.

The checklist, and how to verify each step

1. Count per request, not per day. Split your token logs into per-request numbers: how much of each request is re-read context versus new content. Verify: you can state your median re-read size per request.

2. Separate volume from price, with baselines. Compare request counts week over week, then median context size per request. Volume up means more work; per-request up means the work got heavier. And a rule worth stealing: a fixed cost cannot explain a change, so any "why did X get worse" claim needs the factor's value on a baseline day. Verify: the two trends don't move together. If both are up, you have both problems.

3. A/B your thinking effort. Tag sessions by effort setting and compare replay per request on the same days. Verify: a same-day table like the one above. If the stronger setting shows under ~20% overhead, keep it everywhere and skip this fix.

4. Rank your files by frequency × size × residency. The top of that list is your trim order. Verify: the file you trim first is small and loaded constantly, not big and loaded rarely.

5. Gate every downgrade. For anything leaving the pre-load path, name the scenario that needs it and the mechanism that catches that scenario. Verify: every read-on-demand item has a named catcher. If one doesn't, put it back.

6. Trim your injected index. Measure whatever your client injects into every system prompt, then disable what you don't use weekly. Verify: open a fresh session and compare system-prompt size before and after.

7. Watch one leading indicator. Median re-read context per request, reviewed weekly. Verify: three consecutive weeks flat or down after your fixes. If it creeps back up, something new is pre-loading.

The point under the numbers

Agent economics have the same shape as agent memory: the unit that matters is the single request, not the daily total. Measure per request, attribute against baselines, and put the rule where the mistake happens: a gate that asks "who catches this scenario" at the moment of the change is worth more than any postmortem.

The audit-loop templates referenced in this post are public at wallaby-agent-rules. The companion pieces on memory correctness, why memory needs a closeout ritual and why memory is audited, not written, cover the other half of the same discipline.

Wallaby runs inference for open-weight models with per-token metering and itemized statements, and holds customer bills to the same standard described above. See how we meter or check live status.