Open-weight LLM API pricing, October 2026: Kimi K3, DeepSeek V4, GLM-5.3, Qwen3.8-Max

Every price on this page was fetched from the official pricing page of each provider on October 4, 2026. All prices are USD per 1M tokens. This page is maintained and updated when prices move. The date above is the snapshot, and each number links its source at the bottom.

The one-screen answer, sorted by what a coding-agent workload costs (blended = 70% cached input / 20% fresh input / 10% output, the mix we see in agent loops):

Model (official API) Input Cached input Output Blended (7:2:1)
GLM-5.3-Flash $0.15 $0.03 $0.50 $0.10
DeepSeek V4.1-Flash (peak hours) $0.30 $0.006 $1.20 $0.18
GLM-4.7 $0.60 $0.11 $2.20 $0.42
DeepSeek V4-Pro (peak hours) $1.32 $0.044 $3.96 $0.69
GLM-5.3 $1.40 $0.26 $4.40 $0.90
Qwen3.8-Max (International) $2.00 discount, rate not published $6.00 n/a
Kimi K3 at Wallaby Token $2.70 $0.27 $13.50 $2.08
Kimi K3 (Moonshot official list) $3.00 $0.30 $15.00 $2.31

Disclosure: Wallaby Token sells API access to Kimi K3, and our row is in the table. Every number here is sourced and dated; the sources are listed at the end so you can verify each row yourself.

The five numbers worth quoting

Single lines, each true as of October 4, 2026, each with its source:

  1. Kimi K3's official list price is $3.00 in / $0.30 cached / $15.00 out per 1M tokens, with a 1M-token context window (platform.kimi.ai).
  2. DeepSeek V4.1-Flash costs $0.30 / $1.20 per 1M input/output at peak and exactly half that off-peak; V4-Pro is $1.32 / $3.96 at peak (api-docs.deepseek.com).
  3. GLM-5.3 costs $1.40 in / $0.26 cached / $4.40 out per 1M tokens on z.ai's official API (docs.z.ai).
  4. Qwen3.8-Max costs $2.00 in / $6.00 out per 1M tokens on Alibaba Cloud Model Studio, International deployment (alibabacloud.com).
  5. Wallaby Token sells Kimi K3 at a flat 10% below the official list in every column: $2.70 / $0.27 / $13.50 per 1M (wallabytoken.com/pricing.json).

Four traps these headline numbers hide

1. DeepSeek's price depends on the clock. Peak hours, 01:00–04:00 and 06:00–10:00 UTC on weekdays, cost exactly double the off-peak rate. A workload that runs during Beijing business hours lands squarely in the cheaper window; one that runs during the US evening does not. The $0.18 blended above is the peak rate; off-peak it drops to $0.09. Any page that quotes DeepSeek without naming the window is quoting you half a price.

2. Qwen's older flagship tiers by context length. Qwen3-Max starts at $1.20 in / $6.00 out, but past 128K input tokens per request it climbs to $2.40 / $12.00, and past 256K to $3.00 / $15.00. The newer Qwen3.8-Max is flat to 1M. If your agent reads whole repos, the tier boundary is the price.

3. "Cached input" does not mean the same thing everywhere. Kimi K3's cache-hit tier is 10x cheaper than its miss tier ($0.30 vs $3.00 official). DeepSeek's cache hit is 50x cheaper than its miss. Z.ai lists cached input plus a separate cache-storage line (currently free, marked "limited-time"). For agent workloads that re-read the same context every turn, this cache column decides the bill far more than the headline input price does.

4. Aggregator listings add their own layer. OpenRouter passes provider prices through but charges a fee when you purchase credits. The cheapest K3 listing there today is about $0.99 / $13.00 per 1M from providers with no public track record, well below the official list. Price is one column; uptime history, rate limits and whether the provider will still exist next quarter are the others. We note it because pretending the cheapest listing doesn't exist would make the rest of this page hard to trust.

What one agent task actually costs

Take a coding-agent session that burns 10M tokens at the 7:2:1 mix. Same task, different official APIs:

Model Cost of a 10M-token agent session
DeepSeek V4.1-Flash (peak) $1.84
GLM-5.3 $9.02
Kimi K3 (official list) $23.10
Kimi K3 at Wallaby $20.79
DeepSeek V4-Pro (peak) $6.91

The spread between the cheapest and the most expensive row is about 12x, but these models are not the same product. DeepSeek Flash is a budget tier; K3 is a 1M-context reasoning model. The honest question is not "which row is cheapest" but "which capability tier does your task need, and what does that tier cost."

Where Wallaby sits

We sell one model right now, Kimi K3, so our incentives are simple. Our row:

  • Flat 10% below the official list in every column, including the cache column, which is where agent workloads actually spend.
  • Flat around the clock: the price at 3am is the price at 3pm.
  • Flat to 1M context: the rate at 900K input tokens is the rate at 9K.
  • The listed price is what you pay: no credit-purchase fee on the way in.
  • Every request shows up as an itemized line on your receipt, and our status page is public.

We are not the cheapest row on this page and we don't aim to be. DeepSeek Flash costs less because it is a different class of model. What we sell is the reasoning tier at 10% under list, with billing you can audit line by line. New accounts get $0.50 in free credit, enough for a real test drive.

Methodology and updates

All prices were fetched from the official pages on October 4, 2026 (the Kimi row is also checked daily by our own price monitor). "Blended" assumes a coding-agent mix of 70% cached input, 20% fresh input and 10% output; if your workload is mostly chat, recompute from the raw columns. We update this page when an official price moves, and each update refreshes the snapshot date at the top. Providers run promotions we may not catch on day one, so treat every row as a moment in time.

FAQ

What is the cheapest open-weight LLM API right now? By headline rate, DeepSeek V4.1-Flash at $0.30 / $1.20 per 1M at peak (half that off-peak), and GLM-5.3-Flash at $0.15 / $0.50. Both are budget tiers; at the reasoning tier, Kimi K3's official list is $3.00 / $15.00 and Wallaby sells it at $2.70 / $13.50.

Is Kimi K3 cheaper anywhere than the official price? Yes. Wallaby Token lists K3 at a flat 10% below official in every column. Some aggregator providers list lower still without a public track record behind them; check their uptime and rate limits before routing production traffic.

Do these prices include caching? The cached-input column is listed separately for every model that publishes one. Qwen3.8-Max advertises a context-caching discount but does not publish the rate on its pricing page, so its blended cost is not computed here.

Sources

For provider-level detail on Kimi K3 specifically (17 providers compared), see our earlier snapshot: Kimi K3 pricing compared.