As-of: postmortem covers 2026-09-16 to 2026-09-29 on our production gateway. Every number below comes from our probe archives or billing logs, and the fix described is live.
For thirteen days, we believed the model behind our API had stopped honoring the reasoning_effort parameter. We had evidence: a daily probe showed every effort tier producing suspiciously identical output. We wrote that conclusion into a public pull request and into internal documents. All of it was wrong. The model was fine. Our own router was silently deleting the parameter before requests ever reached the model — and had been, since the day the router went live. This is the postmortem.
The symptom: a ladder that went flat
We run a daily canary that calls the model at every effort tier — none, low, medium, high, xhigh, max, plus parameter unset — and records how much it thinks. On the canary's first day, the tiers were visibly different: the top tier used 344 tokens where medium used 283. From day two onward, the ladder went flat: every tier hit the exact same 256-token ceiling, day after day. Our conclusion: the effort tiers are broken model-side.
Two instrument bugs made that conclusion look better than it was. The canary's own token cap was 256 — a ceiling low enough to flatten the very signal it was measuring; day one had differentiation above the cap, and we didn't read it. And the canary probes through our own gateway rather than the model directly, so any change in our middleware wears the costume of a model-side change.
A change had shipped exactly that day: our failover router went live.
The turn: "works for me"
Thirteen days in, a second look at the model's tier mapping surfaced what we had missed: all six tiers work, none of them disables reasoning — medium behaves as high, xhigh as max, and unset defaults to max — and the tiers only separate if you give the model a token budget far above what our canary allowed. On paper, everything worked.
So we ran the isolation test we should have run first: the same probes, sent directly to the model, bypassing our gateway. Twenty-one calls, a generous token budget, all successful. Under none, the model produced zero reasoning tokens on every run. The tiers were alive. They had been alive the whole time.
The root cause: a whitelist we didn't know existed
Our router calls the model through LiteLLM with drop_params=True — if the model doesn't declare a parameter, drop it silently. For known models, LiteLLM keeps a per-model parameter whitelist. Our model is served under a generic openai/ prefix that this LiteLLM version doesn't recognize, so requests fall back to a generic whitelist — which does not include reasoning_effort. Every call since the router went live had its effort setting quietly deleted, and everything ran at the model's default.
The sharpest detail: LiteLLM's own model tables do list the parameter for gpt-5. One line in a library's lookup table decided whose settings survived.
The fix, and seeing it work
The fix is one line: pass allowed_openai_params=["reasoning_effort"] so the parameter survives the trip. We deployed it the same evening and re-ran the probes through the gateway: none produced zero reasoning tokens, max produced visible thinking. The ladder is back.
We also rebuilt the canary, because a fixed parameter is invisible to a broken instrument: the token cap went from 256 to 2048, the control group became a real none instead of "parameter unset," the verdict now reads reasoning tokens rather than total output, and the test problem is one that actually rewards thinking.
What the retest taught us
With the fix live, we re-ran the ladder on a harder problem — a subset-sum counting task with a known answer — six tiers, three runs each, no token starvation:
| Tier | Median reasoning tokens (3 runs) |
|---|---|
| none | 0 |
| low | 286 |
| medium | 758 |
| high | 426 |
| xhigh | 870 |
| max | 1777 |
Two honest complications fall out of this table:
none is best-effort, not an off-switch. On problems with a clean trick, the model skips thinking entirely — three zeros above. On brute-force arithmetic, it thinks anyway: we measured 4,158 and 8,559 reasoning tokens under none on a harder task. Read none as "please don't," not "cannot."
The middle tiers are not strictly monotonic. Medium out-thought high in this sample (758 vs 426), which matches the mapping table's medium-equals-high auto-promotion. The ends of the dial are solid; the middle has some slack. If you're tuning cost per accepted output, measure your own workload rather than trusting the order of the labels — the price spread across providers matters less than the tier you actually get.
The embarrassing part: what we built on the wrong premise
We had a public pull request on models.dev whose premise was "the tiers don't work." It went through three revisions the night the truth landed — pared back to three defensible tiers aligned with the catalog baseline, with the none wording corrected to best-effort. Internal notes that cited "broken tiers" were corrected, and we confirmed the mapping table was accurate. Correcting yourself in public is cheap. It's not correcting that gets expensive.
What this meant for customers
Between September 16 and September 29, any reasoning_effort setting sent to our gateway was ignored; requests ran at the model's default, which is the maximum tier. If you set a lower tier to control spend, you got — and were billed for — more thinking than you asked for. Every one of those calls is itemized in your usage logs, tokens in and tokens out, per request, so the effect is auditable line by line. The fix has been live since September 29, and the rebuilt canary verifies the ladder daily.
Three rules we're keeping
- Isolate direct vs through-stack first. Before declaring the model broken, run the same probe both ways. Middleware changes wear model-side costumes.
- Your instrument's range is part of the experiment. A 256-token cap turned a working ladder into a flat line. Day-one data held the answer; the instrument hid it.
- Unknown-model fallbacks fail silently. If you proxy models through LiteLLM with
drop_paramsenabled, check what the generic whitelist keeps for your model — before your users do.
Reliability you can verify
We operate a public status page so availability can be verified independently rather than taken on faith — and the canary behind this post watches the gateway the same way, daily. Our terms are published in plain language, and we are a registered Australian company with an ABN on file. The service you use is the service we depend on; this bug annoyed us first.
Get started
Create an account at wallabytoken.com (new accounts receive $0.50 in free credit), mint a key, and set reasoning_effort to whatever your task actually needs — it arrives intact now. Per-request billing settles in seconds, so if you want to watch none behave on your own workload, the itemized logs will show you every token.