Kimi K3 API price per 1M tokens: the full rate card (October 2026)

Three numbers make up the Kimi K3 API rate card: input at $3.00 per million tokens, cached input at $0.30, and output at $15.00. Moonshot AI publishes these as official list, and several providers sell below it. Wallaby's ongoing rate is $2.70 / $0.27 / $13.50. As of October 2026.

Disclosure: Wallaby Token is our own service, so read this as the rate card we can prove, not a neutral review. Every number below links to where it lives.

The rate card

Token type Official list Wallaby
Input $3.00 per 1M tokens $2.70 per 1M
Cached input $0.30 per 1M $0.27 per 1M
Output $15.00 per 1M $13.50 per 1M

Two notes the table cannot hold. Cached input means a repeated prompt prefix served from context cache, billed at one-tenth the input rate. And thinking tokens, the model's reasoning, bill at the output rate, the same rule as the official Kimi API.

What one request actually costs

A coding-agent turn carrying 40,000 input tokens, 36,000 of them cache hits, plus 1,200 output tokens:

  • Official list: $0.0120 + $0.0108 + $0.0180 = $0.0408
  • Wallaby: $0.0108 + $0.0097 + $0.0162 = $0.0367

Four cents per turn. The turn is cheap; the loop is what adds up, and agents loop.

One footnote on the arithmetic: Moonshot also bills a one-time cache-write fee when a prefix is first stored ($3.00 per million written at the 5-minute TTL), and states the write split keeps total cost unchanged from the old all-in input price. Wallaby's September bills reconciled to the three columns above with no separate write line.

Why the cached column decides your bill

Coding agents resend most of their context on every turn, so the majority of input tokens arrive as cache hits at one-tenth price. In one coding-agent session we metered ourselves, 93% of input tokens were cache reads. That is why the cached rate, not the headline input rate, sets the real bill. Mechanics, TTL tiers, and when caching does not help: Kimi K3 cached input pricing.

Providers below official list

A few providers mirror the official $3/$15 exactly, including Together, Fireworks and SiliconFlow. Wallaby sells the same model at a flat 10% below list on every column, ongoing, with per-request itemized billing on top. The full 17-provider table, with cache policies and catches per row, lives in the pillar: Kimi K3 API pricing across 17 providers.

Free options, briefly

The official API has no free tier. Wallaby accounts start with $0.50 of trial credit, enough for roughly 185,000 input tokens. Every other free route has a boundary, and we mapped them honestly: the cheapest way to run Kimi K3.

FAQ

Is Kimi K3 free to use via API?

No. The official API is paid only, and there is no free tier on the rate card. Trial credit is the closest thing: $0.50 on a new Wallaby account, billed per request from the first call.

How much is one token in dollars?

At official list, one input token is $0.000003, one cached input token $0.0000003, one output token $0.000015. Per-token thinking sounds absurd until the loop runs thousands of turns.

Does the price change with context length?

No. The same rates hold up to the 1M-token context window, with no price step at longer contexts.

How often does this price change?

Moonshot sets the list and has changed it rarely. Wallaby's rate tracks at 10% below list. The live source is pricing.json, machine-readable with an as-of date.

Sources

Where to go next

You came for the rate card, so the short version is: $3 / $0.30 / $15 per million at list, $2.70 / $0.27 / $13.50 at Wallaby. Three next steps:

Did this help?