A paper trending this week, SkillRefiner: Offline Skill Refinement from Historical Agent Traces (OpenHands, UT Austin, NYU, CMU; arXiv, October 2026), improves an agent's written skills by mining its past runs: cluster failures and successes separately, derive edits from failure clusters, check them against evidence, and never replays a single task. It reaches 84.0% on SpreadsheetBench, ahead of Trace2Skill at 79.0% and GEPA at 78.5%, while using 1.4–42× fewer refinement tokens.
We know this shape well, because we run a similar loop in production — one we built ourselves, before we knew this paper existed. If you run agents in production, you know the shape of this week: something broke at 2 AM, the fix shipped, and nobody went back to ask whether the fix actually worked. Our operating loop exists to make that question unskippable. Here it is, with the receipts.
The pipeline, five stages
Every failure our agent makes becomes a regression case at the moment it happens, with a written re-test condition rather than a bare description. A mechanical scan clusters open cases by root cause; two or more recurrences in a family is the threshold for action, and a single occurrence is treated as noise. A fix can only be marked verified with evidence attached: the verification method written at registration must have actually been observed. Then a weekly post-evaluation pass asks two questions of every open case: did the same family come back, and was the fix complete. What survives becomes a one-line red line inside the agent's standing instructions — scoped to the context where the failure happened, not applied globally.
You can walk the pipeline below: every card opens a real case from our log, with supplier names, customers, and pricing removed:
Pick a step above.
Three receipts
The evidence gate stopped a wrong-layer fix. A monitoring probe began failing with HTTP 403, and the alert text blamed an external provider. The actual cause was $0.000278 remaining on the probe's own key: the gateway was rejecting the probe, while the external provider had hundreds of dollars in balance. Because our failure entry records the exact error signature, the durable fix went where it belonged: the alert text now maps "403 + quota wording" to "probe key exhausted, top up this exact row." Total customer impact: zero. But without the signature, we would have "fixed" a healthy provider.
Clustering ended a 13-day misdiagnosis. For thirteen days we believed the model ignored a reasoning-effort parameter. A controlled A/B test of the same request, direct versus through our routing layer, proved our own middleware was silently dropping the parameter. (We wrote that incident up separately: why reasoning_effort silently stopped working.) That case joined a family of "the middleware ate the signal" incidents. The paper's framing matches how we treat it: singletons are anecdotes, families are architecture bugs. Because this was a family, the fix went into the router for the whole family rather than one ticket.
Post-evaluation caught a fix that was only documentation. One early fix read, in effect, "be more careful when reporting status." The same miss recurred eight days later: the fix had changed words, not behavior. It was replaced with a mechanical consistency scan that catches the divergence by itself. The same logic is why a 'stale memory' case stayed open for two review cycles until three consecutive weekly audits showed zero recurrence. Closure is earned, not declared.
What the paper taught us back
Two things, immediately adopted.
Scope the fix to the failure's context. SkillRefiner's merge step restricts a behavior only where the failure was observed, keeping it everywhere else. Our instinct after an incident is to add a global rule; one incident, one new global law, forever. Our failure-entry format now requires every fix to name its scope ("in this specific context"), because unscoped rules are how instruction files bloat into contradiction.
Failure cases outweigh success cases, by a lot. The paper's ablation removes failure-driven edits and loses 8.0 accuracy points; removing success-driven reinforcement changes nothing (±0.0). That matched our gut so exactly it was almost annoying. Our negative casebook gets cited in drafting sessions weekly; our positive one is mostly a trophy shelf. We keep both, but the weighting is no longer a gut call.
The loop is the point
None of this is clever. A log, a threshold of two, an evidence field, a weekly review, and the discipline to write the lesson where the agent actually reads it. What the paper adds is proof that this loop, run offline from history with no replays, is not housekeeping. It is the mechanism that makes the agent next month better than the agent this month. And because we audit the memory itself, the loop is inspectable end to end.
We will keep running ours. The paper is welcome to keep explaining why it works.