Where to get Kimi K3: every access route compared

There are six ways to run Kimi K3 today, and the right one depends on your billing model, not on whether K3 is available. It is, almost everywhere. This page gives you the full route table, current prices per route, and a one-line pick for each situation. Reading time: about 6 minutes. If you just want the answer: most individual developers should start on a free tier or a $0.50 trial, most teams should pick a flat per-token provider with itemized billing, and almost nobody should self-host a 2.8T-parameter model.

Kimi K3 is Moonshot AI's open-weight reasoning model: 1M-token context, native vision, strong coding and agentic benchmarks. The weights are public, so access routes multiply fast. Here is the map as of early October 2026.

Disclosure: Wallaby Token sells API access to kimi-k3, and our route is in the table. Prices below come from each provider's public page, dated; sources are listed at the end.

The routes at a glance

Route Example providers Price signal, USD per 1M tokens Best fit
Official API Moonshot platform.kimi.ai $3.00 in / $0.30 cached / $15.00 out Day-one access to new checkpoints, reference behavior
Flat per-token providers Wallaby Token, SiliconFlow, Together, Fireworks $2.70–$3.00 in / $13.50–$15.00 out Predictable bills, OpenAI-compatible clients
Gateways OpenRouter, Vercel AI Gateway $0.88–$3.45 in / $7.86–$17.25 out, 19 endpoints One key across many models, failover
Discount hosts SayGM, inference.net, DeepInfra $1.03–$2.85 in / $5.17–$14.25 out Price-first workloads you can re-verify
Subscriptions GitHub Copilot Pro, OpenCode Go $10/month flat Agents inside one tool, light monthly volume
Self-host Hugging Face weights ~$2M hardware or $15–25K/month cloud GPUs Research and fleet-scale volumes only

Official API: the reference route

Moonshot's own platform lists $3.00 input, $0.30 cached input, $15.00 output, and it is where new checkpoints land first. Two caveats show up in third-party measurements: Artificial Analysis has clocked Moonshot's endpoints as slower than rival hosts on earlier Kimi models, and requests are processed in Beijing, which some security teams will not sign off on regardless of policy text. If neither matters to you, direct is the cleanest line.

Flat per-token providers: the predictable route

Several providers, including Together, Fireworks and SiliconFlow, mirror the official $3/$15 list exactly. A few undercut it: Wallaby's launch pricing is $2.70 / $0.27 / $13.50 through October 31, 2026, a straight 10% off in every column. What separates this group is not the headline rate, it is the billing shape: per-request itemized statements, per-key budget caps, and a public status page are the difference between a provider you can hand to a team and one you cannot. Fireworks adds a US-hosted, zero-retention variant at roughly a 10% premium; SiliconFlow speaks both OpenAI and Anthropic request formats.

Gateways: one key, many models

OpenRouter listed 19 Kimi K3 endpoints as of late September, with input prices from $0.88 to $3.45 and output from $7.86 to $17.25. Read the cheap rows carefully: the lowest prices are fp4 or fp8 quantized endpoints, while Moonshot's own endpoint is native mxfp4 at list price. Test quality before optimizing on price. The gateway pitch is consolidation: one key, one bill, failover across hosts. You pay for it in a fee on top of the endpoint rate.

Discount hosts: cheap, but verify

SayGM lists K3 at $1.03 input / $5.17 output, including a TEE (trusted execution environment) route at the same price. inference.net sits at $2.10 / $10.95, DeepInfra and Phala at $2.85 / $14.25. These are real prices on public pages, and they are below what the model's maker charges. Sustained sub-list pricing usually means subsidized capacity or quantized serving; neither is automatically bad, but treat discount endpoints as re-verifiable infrastructure, not permanent fixtures.

Subscriptions and free tiers: the evaluation route

For a first look, free is honest: NVIDIA and Zenmux serve K3 at $0 with rate limits, and Fireworks' Fire Pass has featured K3 on its free tier. For agent-first evaluation, OpenCode Go runs $5 for the first month, then $10. GitHub Copilot Pro at $10/month includes K3 in its model picker. Subscription math flips quickly: at metered rates, a 3:1 input:output blend of K3 costs about $6 per million tokens, so a $10 subscription only wins below roughly 1.7M tokens a month.

Self-hosting: the $2M route

The weights are on Hugging Face under the Kimi K3 License. A single production replica needs on the order of 8× H100-class GPUs, around $2M in hardware or $15–25K per month rented. Break-even against hosted API access sits near 50 million tokens per day. Below that volume, self-hosting is a research project, not a cost strategy.

Which route should you pick?

  • Just looking: a free tier from NVIDIA, Zenmux or Fire Pass, or a $0.50 trial credit; zero commitment either way.
  • Individual developer, coding agent: a flat per-token provider with a trial; you will know your real cost within an afternoon.
  • Team lead: flat per-token with per-key budget caps and itemized statements. The bill shape matters more than a 10% rate difference.
  • Already on a gateway: stay, but pin a specific endpoint rather than "cheapest" routing if quality matters.
  • Regulated or US-only: Fireworks' US-hosted variant, or any TEE route, after reading the attestation fine print.
  • 50M+ tokens/day, every day: now you may price self-hosting.

One trap regardless of route: the cache column. Agent workloads resend most of their context every turn, so cached-input price decides real bills more than headline input price. We measured this in Kimi K3 pricing compared: official list vs 17 providers: the blended spread between rows reaches 6× on the same workload. And if you are comparing providers on price alone, measure cost per accepted output instead: retries and failed generations are the bigger bill.

FAQ

What is the cheapest way to try Kimi K3? A free tier from NVIDIA, Zenmux, or Fireworks' Fire Pass if you only want a look, or a provider with a small trial credit if you want to test a real workflow. Both cost nothing upfront.

Is a subscription or per-token billing cheaper for Kimi K3? Below roughly 1.7M tokens a month, a $10 subscription wins on pure price. Above that, metered billing at ~$6 per million blended tokens wins, and metered billing is the only shape that scales honestly for a team.

Are discount Kimi K3 endpoints the real model? Sometimes. Quantized fp4/fp8 endpoints are cheaper to serve and price lower; fingerprint checks and a fixed test prompt suite are how you verify. Re-check after any price change.

What you built, and where to go next

You now have the full route map and a pick for your situation. Three useful next steps:

  • Deeper: the API docs: quickstart is four steps, and new accounts get $0.50 of trial credit.
  • Sideways: Open-weight LLM API pricing, October 2026: K3 against DeepSeek V4, GLM and Qwen on the same yardstick.
  • Forward: pricing: the live Wallaby rate card, machine-readable at /pricing.json.

Sources: Moonshot platform pricing page; OpenRouter endpoint list (September 28, 2026); SayGM, inference.net, DeepInfra, Phala, SiliconFlow, Together, Fireworks public pricing pages (September–October 2026); Artificial Analysis provider measurements; llmprice subscription data (September 28, 2026); Eden AI self-host cost estimate (August 2026). All prices USD per 1M tokens unless noted; verify against the provider's page before committing budget.