Kimi K3 cached input pricing: $0.30 per million, and when caching pays (October 2026)

A cache hit on Kimi K3 costs $0.30 per million input tokens, one tenth of the $3.00 miss price. Agent workloads that resend the same context every turn live on this column: at a 93% hit rate, the blended input price falls to about $0.49 per million, an 84% cut. As of October 2026.

Disclosure: Wallaby Token is our own service. The official numbers below come from Moonshot's published documentation, linked where they live; the session math comes from our own metered runs.

The cache rate card

Moonshot splits caching into three billing items, per the official context-caching docs:

Billing item Price per 1M tokens What it is
Input (cache miss) $3.00 The portion of a request that misses the cache
Cache Write, 5m TTL $3.00 Charged once when a prefix is first written; each hit renews it free
Cache Write, 1h TTL $6.00 Same, with a one-hour lifetime
Cached input (cache hit) $0.30 The portion served from cache, per request

Every hit saves $2.70 per million tokens against the miss price. On Wallaby the hit itself is cheaper too: cached input bills at $0.27 per million, 10% under official list, per pricing.json.

Whose fee is the write fee? Moonshot's. On Wallaby, September's metered bills reconciled to the three columns above — input, cached input, output — to the cent, with no separate write line. Moonshot, for its part, states the write split keeps total cost unchanged from the old all-in input price.

How the cache actually works

Caching matches request prefixes. Stable content goes at the front (system prompt, tool definitions, codebase), changing content at the end; if any part of a prefix changes, everything after that position misses. Entries live 5 minutes by default, or 1 hour if you set prompt_cache_options.ttl, and a hit renews the entry under the original TTL at no charge. Cache is isolated per organization, and entries cannot be cleared manually.

One structural caveat: cache is stored in blocks, so a prefix portion smaller than one full block is billed as a miss regardless. In practice, a small edit inside a block can turn an expected hit into a miss for a span much larger than the edit itself.

The math on a real session

In a coding-agent task we metered ourselves (five requests, Codex), 85k of 91k input tokens came back as cache reads. Input-side cost, official list: $0.27 without caching, $0.044 with it. That single knob did more for the bill than any provider discount could.

At a 93% hit rate the blended input price is 0.93 × $0.30 + 0.07 × $3.00 ≈ $0.49 per million, versus $3.00 uncached. The hit rate is the whole game, and it is set by your client's behavior, not your provider's generosity.

When caching does not help

Honest boundaries, because they exist. Single-shot short sessions have no repeated prefix to reuse. If your interval between calls exceeds an hour, entries expire before they pay off. And low hit rates are not automatically a bug: in our own September bill, release-tracking jobs ran at a 12% hit rate because each run scans fresh changelogs, and that shape is correct for the job. Judge hit rates by call shape, not by a universal benchmark.

5m or 1h

The official math, condensed: choosing 1h over 5m costs exactly $3.00 more per million written, and each hit saves $2.70, so two hits in the hour put you ahead. Agents in continuous loops fit the 5m default; long sessions with human pauses fit 1h. The TTL locks at first write, so switching means waiting for the entry to expire.

FAQ

What is the cached input price for Kimi K3?

$0.30 per million tokens at official list, $0.27 on Wallaby. One tenth of the miss price either way.

How long does the cache last?

Five minutes by default, one hour with prompt_cache_options.ttl: "1h". Every hit renews the entry for free under the original TTL.

Do all providers charge the same cache rate?

No. The cache column varies more than the headline input price across providers, which is why the 17-provider comparison carries a cache column at all.

How do I know if my requests are hitting cache?

The API returns it per request: usage.prompt_tokens_details.cached_tokens on Chat Completions. A minimal response tail looks like this:

"usage": {
  "prompt_tokens": 10240,
  "prompt_tokens_details": { "cached_tokens": 9216 },
  "completion_tokens": 512
}

On Wallaby, every request lands as an itemized line in your console with cached input metered separately.

Does Wallaby charge cache-write fees?

Our bills itemize three columns: input, cached input, output. September's metered spend reconciled to those three exactly. The write-fee rows above are Moonshot's official billing items, quoted so you can read their documentation without surprises.

Sources

Where to go next

You came for the cache price, so the short version is: hits cost $0.30 per million, misses $3.00, and your hit rate decides which one you mostly pay. Three next steps:

Did this help?