What our own agent stack costs on Kimi K3: a real September bill (2026)

We run our own agent stack on Kimi K3 through our own gateway: task management, availability testing, tutorial verification, release tracking, benchmarks. In September 2026 it burned about 4.2 million tokens and cost us $17.69, metered per request. Here is the bill, broken down by task.

Disclosure: Wallaby Token is our own service, and this is our own metered usage on it. Every figure below comes from our gateway's request log, not from list-price arithmetic.

The September bill

Task Requests Input tokens Output tokens Cache hit rate Cost
Task management 326 568K 391K 12% $6.66
Availability testing 1,650 204K 448K 52% $6.35
Tutorial verification 141 2,210K 65K 72% $3.01
Release tracking 847 120K 48K 12% $0.95
Benchmark runs 180 57K 35K 50% $0.55
Everything else 33 40K 5K 4% $0.17
Total 3,177 3.20M 0.99M 56% $17.69

Hit rate here is the share of input-token volume served from cache, not the share of requests. Rates are our published ones: $2.70 / $0.27 / $13.50 per million for input, cached input and output. Same model, same month, wildly different cost shapes.

What each row teaches

Output is the wallet-buster. Availability testing sent only 204K input tokens all month and still spent $6.35, because 95% of that was output: the model writes long evaluations, and output costs five times input. Task management tells the same story at 79% of spend. If you are trimming a bill, trim generated tokens first.

Caching rescues input-heavy work. Tutorial verification pushed 2.2M input tokens through the model, the most of any row, yet cost $3.01. 72% of that input came back as cache reads at one-tenth price. Without caching, at our rates, that row alone would have cost $6.85 instead of $3.01.

Hit rate follows call shape, not virtue. Release tracking scans fresh changelogs each run, so a 12% hit rate is correct for the job. Benchmarks re-send identical prompts, so 50% comes free. There is no universally good number.

Two coding sessions we metered

The rows above are automation. For developer-shaped work, two sessions we metered ourselves: a five-request Codex task totalled $0.16, with 93% of its input served from cache; four Cline runs of the same task ranged from $0.42 to $1.43 depending on reasoning effort. Single sessions, small samples, honestly labelled. They say a solo developer's daily driving lands in single-digit dollars per month more often than not.

Estimating your own month

Take your monthly volumes and apply:

cost = (uncached_input × $2.70 + cached_input × $0.27 + output × $13.50) / 1,000,000

Our September bill reconciles to this formula to the cent. (Moonshot's official pricing additionally lists a one-time cache-write fee per prefix written; the cache pricing post has that table.)

If you do not know your cache hit rate yet, run one typical day and read cached_tokens off the usage field. The estimate you get from that one day of real numbers beats any table we could publish, including this one.

FAQ

How much does Kimi K3 cost for light use?

Single-digit dollars per month. Our benchmark and misc rows, together 213 requests, cost $0.72 combined. Testing costs less: the $0.50 trial credit covers about 50 smoke-test calls.

Is there a monthly subscription for the API?

No. It is pay-as-you-go: prepaid top-ups from $20, the balance never expires, and rates stay identical on trial credit and paid balance.

What does a typical coding session cost?

In our metered sessions, between $0.16 and $1.43 depending on task size and reasoning effort. Cache hit rate is the biggest lever inside that range.

How do I cap my monthly spend?

On Wallaby, create a named key per developer or per pipeline, each with its own dollar cap. When the cap is hit, calls stop. Every request lands as an itemized line, so the cap never surprises you.

Sources

Where to go next

You came for a monthly number, so the short version is: our whole agent stack costs us $17.69 a month, and your number depends on output volume and cache hit rate more than anything else. Three next steps:

Did this help?