Kimi K3's official list price is $3.00 per million input tokens, $0.30 per million cached-input tokens, and $15.00 per million output tokens. That is the number Moonshot AI publishes — and ten of the seventeen providers below match it on headline input and output. The rest charge less, sometimes much less: on a coding-agent workload, the spread between the cheapest and the most expensive paid row is about 6×.
The one-screen answer (all prices USD per 1M tokens; blended assumes a coding-agent mix of 70% cached input / 20% fresh input / 10% output):
| Input | Cached input | Output | Blended (7:2:1) | |
|---|---|---|---|---|
| Moonshot AI (official list) | $3.00 | $0.30 | $15.00 | $2.31 |
| Wallaby Token | $2.70 | $0.27 | $13.50 | $2.08 |
| Cheapest paid row in this snapshot (Crof, eco endpoint) | $1.00 | $0.10 | $4.00 | $0.67 |
Free tiers exist — NVIDIA and Zenmux serve Kimi K3 at $0 with rate limits, fine for a first look, not for a workload. And our rate is a straight 10% off the official list in every column, including the cache column, which is where agent workloads actually spend.
Everything below is the long answer for Kimi K3, Moonshot AI's open-weight reasoning model: the full 17-provider table, why the cache column decides real bills, how the OpenRouter row should be read, and what a finished task actually costs.
Disclosure: Wallaby Token sells API access to kimi-k3, and our row is in the table. Every number below is sourced and dated; the sources are listed at the end, and you can verify each row yourself.
If your workload is mostly chat rather than agent loops, your mix will differ from the 7:2:1 blend — the raw columns are there so you can recompute.
The table (accessed 2026-09-13)
Snapshot date: 2026-09-13 (updated 2026-09-14 to add Alibaba Cloud Model Studio). Several providers run rotating discounts — OpenRouter's listed price moved twice within days of this snapshot — so treat every row as a moment in time and re-check the ones you care about.
| Provider | Input | Cached input | Output | Blended (7:2:1) | Notes |
|---|---|---|---|---|---|
| NVIDIA (build.nvidia.com) | free | free | free | $0 | Rate-limited evaluation tier, not for production |
| Zenmux (free endpoint) | free | free | free | $0 | Free tier with limits; paid tier at list price |
| Crof (eco endpoint) | $1.00 | $0.10 | $4.00 | $0.67 | 1M context |
| Crof | $2.00 | $0.25 | $8.00 | $1.38 | 1M context |
| Nano-GPT | $2.00 | $0.20 | $10.00 | $1.54 | 1M input limit; also a TEE variant |
| OpenRouter | $2.65 | $0.30 | $13.28 | $2.07 | Routes across upstreams; runs rotating discounts (5–15% off observed after this snapshot), price drifts weekly |
| Wallaby Token | $2.70 | $0.27 | $13.50 | $2.08 | 10% off the official list in every column, ongoing |
| Alibaba Cloud Model Studio (Global) | $2.827 | $0.283 | $14.133 | $2.18 | Global-scope list price |
| DeepInfra | $2.85 | $0.285 | $14.25 | $2.19 | Image input supported |
| Moonshot AI (official) | $3.00 | $0.30 | $15.00 | $2.31 | The reference list price |
| Together AI | $3.00 | $0.30 | $15.00 | $2.31 | |
| Fireworks AI | $3.00 | $0.30 | $15.00 | $2.31 | Image input supported |
| Novita AI | $3.00 | $0.30 | $15.00 | $2.31 | 1M context |
| Vercel AI Gateway | $3.00 | $0.30 | $15.00 | $2.31 | 1M context |
| Chutes (TEE) | $3.00 | $0.30 | $15.00 | $2.31 | Output capped at 65k tokens |
| Requesty | $3.00 | $0.45 | $15.00 | $2.42 | Output capped at 262k |
| Synthetic | $3.00 | $0.45 | $15.00 | $2.42 | 512k context |
| Baseten | $3.00 | no cache price listed | $15.00 | ~$4.20 | Without a cache discount, cached tokens bill as input |
| Nebius | $3.00 | $3.00 | $15.00 | $4.20 | Cache billed at full input price |
Free tiers (NVIDIA, Zenmux) are legitimate for evaluation but rate-limited — useful for a first look, not a workload.
What the table actually tells you
1. The headline input price is nearly uniform — the differences are in cache and output. Ten providers sit at exactly the official $3.00 / $15.00 list price. What separates a $2.31 bill from a $4.20 bill on the same workload is the cached-input price. Baseten and Nebius have the same headline as everyone else but bill cache reads at full price, which nearly doubles the cost of an agent workload.
2. Cache price matters more than any other column if you run coding agents. Agent frameworks (Codex CLI, Cline, Claude Code) resend a large system prompt and conversation history on every turn — in our Codex run, 85k of 91k input tokens were cache reads. At a 7:2:1 mix, a provider's cache column carries 70% of your traffic. A "cheap" provider with no cache discount is not cheap.
3. Output limits differ and they bite silently. Chutes caps output at 65k tokens; kimi-k3 is a reasoning model that can think at length, so a cap truncates long reasoning chains. If your tasks need long outputs, check the limit column before the price column.
4. The cheapest rows are cheap for a reason. Sub-$1.50 blended prices are real, and for many workloads those providers are a good choice. What you give up varies by provider: rate limits, TEE-only endpoints, fewer regions, or pricing that changes without a published policy. Check the notes, test with your actual workload, and keep a second provider configured.
5. Comparing against OpenRouter? The rows are a cent apart — read the notes column. At this snapshot, OpenRouter's blended price ($2.07) and our rate ($2.08) differ by one cent. What the number alone doesn't show: OpenRouter routes across upstream providers, so the upstream serving your request can change underneath you, and it runs rotating discounts — its listed K3 price moved twice within days of this snapshot. Treat that row as a moment in time and re-check it before budgeting. The same discipline applies to every row, ours included.
A worked example with real tokens
Theory aside, here is a real coding-agent task — the bookstore landing page from our Codex CLI guide: 5 requests, 91,155 input tokens (85k from cache), 5,987 output tokens.
| Wallaby | Official list price | A no-cache-discount provider | |
|---|---|---|---|
| Cached input (85k) | $0.0230 | $0.0255 | $0.2552 |
| Fresh input (6k) | $0.0165 | $0.0183 | $0.0183 |
| Output (6k) | $0.0808 | $0.0898 | $0.0898 |
| Total | $0.12 | $0.13 | $0.36 |
Same model, same task, same tokens. Two honest ways to read this table: against the official list price — the comparison most buyers actually face — our rate saves about 8% on this workload. Against a provider with no cache discount — the worst case, included here because two major providers price this way — the spread is 3×, driven almost entirely by the cache column. (Figures computed from the published per-token prices above; the actual itemized bill for the Wallaby column is shown in the Codex guide.)
The price column is not the bill
Everything above answers "what does a token cost." The question that actually decides budgets is "what does a finished task cost" — retries and human fixes appear on no pricing page. The honest unit is cost per accepted output: cost per attempt ÷ acceptance rate. A cheaper token with a lower acceptance rate can lose to a pricier one that ships on the first try. Measuring yours takes an afternoon and twenty real tasks; the method is in our cost-per-accepted-output playbook.
What to check beyond price
Price is one column. And per-token price is only the visible cost — rework is the bigger one, which is why cost per accepted output beats price lists as a comparison unit. Before committing a workload, we suggest verifying five things about any provider — including us:
- Machine-readable pricing, public. A pricing.json or equivalent you can poll without an account. If you have to log in — or email someone — to learn the price, budget planning is guesswork.
- A trial path measured in minutes. A small free credit and a self-serve key beat a sales call for evaluation.
- Itemized billing. Every request should appear with its token counts, so you can reconcile spend against your own logs.
- A public status page. So you can verify availability independently before debugging your own config.
- Model policy. Which models get added and why. We host open-weight models only, and add them based on sustained usefulness rather than launch-cycle attention.
Sources and method
- Provider prices: models.dev open data repository, provider configuration files accessed 2026-09-13, cross-checked against provider pricing pages where public. Alibaba Cloud Model Studio row added 2026-09-14 from its international-site pricing documentation.
- Blended weighting: Artificial Analysis blended-price methodology (70% cached input / 20% input / 10% output).
- Wallaby Token prices: our own pricing.json, live.
- Worked-example token counts: our own billing dashboard, reproduced in the Codex guide.
Prices change. Check the date at the top of the table, and verify the rows you care about before making a decision.
Still choosing between routes rather than rows? Where to get Kimi K3: every access route compared maps official API, flat per-token providers, gateways, discount hosts, subscriptions and self-hosting against each other. Comparing across models instead of providers? Our open-weight LLM pricing snapshot lines up Kimi K3, DeepSeek V4, GLM and Qwen on one table. And if the deciding factor is the checkout itself, Kimi K3 with card payment and no crypto walks through paying by card, itemised invoices and refunds. For our own track record, the public status page shows per-component uptime over 90 days and a machine-readable feed.
Wallaby Token is an inference platform for open-weight models. New accounts get $0.50 in trial credit — enough to run the worked example above several times over. Sign up · Live pricing · Status