Three numbers make up the Kimi K3 API rate card: input at $3.00 per million tokens, cached input at $0.30, and output at $15.00. Moonshot AI publishes these as official list, and several providers sell below it. Wallaby's ongoing rate is $2.70 / $0.27 / $13.50. As of October 2026.
Disclosure: Wallaby Token is our own service, so read this as the rate card we can prove, not a neutral review. Every number below links to where it lives.
The rate card
| Token type | Official list | Wallaby |
|---|---|---|
| Input | $3.00 per 1M tokens | $2.70 per 1M |
| Cached input | $0.30 per 1M | $0.27 per 1M |
| Output | $15.00 per 1M | $13.50 per 1M |
Two notes the table cannot hold. Cached input means a repeated prompt prefix served from context cache, billed at one-tenth the input rate. And thinking tokens, the model's reasoning, bill at the output rate, the same rule as the official Kimi API.
What one request actually costs
A coding-agent turn carrying 40,000 input tokens, 36,000 of them cache hits, plus 1,200 output tokens:
- Official list: $0.0120 + $0.0108 + $0.0180 = $0.0408
- Wallaby: $0.0108 + $0.0097 + $0.0162 = $0.0367
Four cents per turn. The turn is cheap; the loop is what adds up, and agents loop.
One footnote on the arithmetic: Moonshot also bills a one-time cache-write fee when a prefix is first stored ($3.00 per million written at the 5-minute TTL), and states the write split keeps total cost unchanged from the old all-in input price. Wallaby's September bills reconciled to the three columns above with no separate write line.
Why the cached column decides your bill
Coding agents resend most of their context on every turn, so the majority of input tokens arrive as cache hits at one-tenth price. In one coding-agent session we metered ourselves, 93% of input tokens were cache reads. That is why the cached rate, not the headline input rate, sets the real bill. Mechanics, TTL tiers, and when caching does not help: Kimi K3 cached input pricing.
Providers below official list
A few providers mirror the official $3/$15 exactly, including Together, Fireworks and SiliconFlow. Wallaby sells the same model at a flat 10% below list on every column, ongoing, with per-request itemized billing on top. The full 17-provider table, with cache policies and catches per row, lives in the pillar: Kimi K3 API pricing across 17 providers.
Free options, briefly
The official API has no free tier. Wallaby accounts start with $0.50 of trial credit, enough for roughly 185,000 input tokens. Every other free route has a boundary, and we mapped them honestly: the cheapest way to run Kimi K3.
FAQ
Is Kimi K3 free to use via API?
No. The official API is paid only, and there is no free tier on the rate card. Trial credit is the closest thing: $0.50 on a new Wallaby account, billed per request from the first call.
How much is one token in dollars?
At official list, one input token is $0.000003, one cached input token $0.0000003, one output token $0.000015. Per-token thinking sounds absurd until the loop runs thousands of turns.
Does the price change with context length?
No. The same rates hold up to the 1M-token context window, with no price step at longer contexts.
How often does this price change?
Moonshot sets the list and has changed it rarely. Wallaby's rate tracks at 10% below list. The live source is pricing.json, machine-readable with an as-of date.
Sources
- pricing.json: Wallaby rates, live, asOf 2026-09-08
- platform.kimi.ai pricing docs: official list and caching rules
- models.dev: third-party price index
Where to go next
You came for the rate card, so the short version is: $3 / $0.30 / $15 per million at list, $2.70 / $0.27 / $13.50 at Wallaby. Three next steps:
- Wider: Kimi K3 API pricing across 17 providers — the full comparison table.
- Deeper: Kimi K3 cached input pricing — the column that decides your bill.
- Cheaper: the cheapest way to run Kimi K3 — paid floors, free routes, catches.