A new Microsoft Research paper, Agensh, ran 1,024 AI agents on a single task with no manager and watched them organize themselves. The headline result is real: on one brutal task, the pass rate climbed from 33.89% with one agent to 55.06% with 1,024. But if you came here asking what a run like that costs — nobody knows. The paper never says. No token counts, no dollar figures, no cost per point gained. That silence is the story, because agent count has quietly become a new scaling axis for AI capability — and every scaling axis is also an invoice axis.
What Microsoft actually ran
The setup is genuinely novel, so a quick sketch. Most multi-agent systems are orchestrator-worker: a boss agent decomposes the task and coordinates the team, and the boss's capacity becomes the ceiling. Agensh removes the boss. Every one of the 1,024 workers gets the identical prompt (only the ID differs) and self-organizes through three pieces of shared infrastructure: a git server holding the work, a chat system for announcements and direct messages, and an append-only context board where agents post typed short entries — facts, failed hypotheses, claims on subtasks, patch summaries.
The workload was ProgramBench: five of its hardest tasks, each a six-hour, offline, from-scratch rebuild of a real software project (pandoc, PHP-src, FFmpeg, and friends, tens of thousands to millions of lines), powered by a frontier reasoning model.
As agent count grew, the researchers watched organizational behavior emerge in layers: at 8 agents, peers negotiated interfaces; at 32, they reviewed and revoked each other's work; at 128, they standardized processes; at 1,024, they settled into org-level roles, with multiple agents serving as "integrators" and requesters shopping among them. Whatever your priors on multi-agent hype, this is a serious experiment and the emergence observations are worth the read.
The curve nobody plotted
Here is the result everyone quoted, and the shape nobody drew:

Two ways to read that curve:
- The generous reading: going from 1 to 128 agents lifted the five-task average pass rate from 19.31% to 28.78% — a relative gain of about 49%, from nothing but adding peers.
- The invoice reading: on pandoc, the one task pushed to 1,024 agents, the jump from 128 to 1,024 agents (an 8× multiplication of the workforce) bought +4.1 percentage points. The first 128 agents bought +16.9. Each successive doubling of headcount is worth less than the last, and the curve is bending hard.
This is not a criticism of the paper; diminishing returns are what honest scaling curves look like. It is a criticism of the coverage, which quoted "+21 points" and moved on. A result like this has two axes: what you gained and what you spent. Only one of them made the headlines.
The missing line item
Which brings us to the number that isn't there. Nowhere in thirteen pages does the paper report token consumption, dollar cost, or marginal cost per percentage point. We know the shape of the bill: 1,024 agents, six hours, a frontier reasoning model, plus the coordination overhead of a shared context board every agent reads. Beyond that, nothing.
That omission matters more than it might seem, because the paper's own protocol quietly agrees that budget is the binding constraint. The highest-value entry type on the shared context board, the authors write, is the FAIL entry, a falsified hypothesis, precisely because "it stops peers spending budget on it." Even in a lab with Microsoft-scale resources, the point of the whole coordination layer is to stop a thousand agents from burning budget on dead ends. Cost isn't a footnote to multi-agent systems. It is the thing the coordination is for.
And if this pattern catches on — and "agent count as a scaling axis" is exactly the kind of result that catches on — workloads change shape. A task that used to mean one agent making a few dozen model calls becomes an organization of hundreds making calls for hours. Multi-agent runs multiply token consumption per task by orders of magnitude compared to a single agent. For the people paying, that's not an architecture detail. It's the budget.
If you're the one paying
Three practical notes, from the side of the counter that reads bills:
- Agent count is now a knob with a marginal price. Like model size or reasoning effort, it's a dial you turn for capability — and like every dial, it only saves or costs money if it actually reaches the model and you measure its effect. (Ask us how we know.) The right question is never "how many agents can we run" but "what does the next agent buy us on this task?"
- Measure at the task level, not the token level. A swarm that finishes the job beats a solo agent that doesn't, regardless of the per-token price. And a swarm that adds 4 points for 8× the spend is a bad trade at any price. The honest unit is cost per accepted output: total dollars divided by results you'd actually ship. Multi-agent setups make that measurement more important, not less.
- Scale headcount only where the task pays for it. The Agensh curve is sublinear, so the economic sweet spot is likely far below the headline number for most workloads. Start with the fewest agents that finish the task; add more only when the marginal points are worth their invoice.
The honest fine print
Worth saying, because the paper itself is careful and the coverage hasn't been: Agensh ran no controlled comparison against an orchestrator-worker system at equal budget, so "no boss" is an architecture that works, not a proven superior one. The experiments are one task family, one model, one harness. And 55% is a pass rate, not a solved task. What the paper demonstrates is that self-organization works at a scale nobody had shown before. What it costs, and whether it's worth it, is exactly the work that hasn't been done yet.
Somebody should do that work. We'll be reading.
Wallaby Token is an inference platform for open-weight models — the kind of meter a 1,024-agent run would spin hard, which is why we read papers for the line items. We publish our prices in the open at wallabytoken.com/pricing.json, and we'd rather be compared on cost per accepted output than on token lists.