Dispatch / tasks
Methods & practice
How we run memory, audits, and closeout rituals in production.
15 notes · recently updated first
- 2026-10-06 AI Memory Should Have a Probation Period Trust-building mode: new memories get proposed, not written, until you approve them. Plus a conditional code-version module and a full Chinese edition. Playbook
- 2026-10-05 AI Memory Benchmarks Measure the Wrong Thing Every major AI memory system posts green benchmark scores, and every agent user reports the same failures. The gap is structural: benchmarks test retrieval, while production memory dies of rot. Playbook
- 2026-10-05 Why your agent's token bill keeps growing, and how to audit it An agent's daily token spend can double with no change in workload. The causes are measurable once you count per request instead of per day. A field guide with real numbers and a cost simulator. Playbook
- 2026-10-04 SkillRefiner: five stages, zero replays — like ours SkillRefiner refines agent skills offline from historical traces. We run a similar loop in production, built independently: five stages, real cases, and what it taught us back. Playbook
- 2026-10-01 Your Agent Needs Sleep, Not a Bigger Context A stale card on our project board exposed the real bug: we had built our AI a perfect memory, and perfect memory was the problem. Human memory works because it forgets — so we rebuilt ours around sleep, pruning, and myelin. Playbook
- 2026-10-01 Memory Consolidation for AI Agents: A Three-Piece Playbook A working recipe for agent memory that gets more accurate over time: a replay job that decides what survives the day, a pruning pass that deletes on purpose, and a reflex layer that turns repeated procedures into muscle. Playbook
- 2026-10-01 Memory Is Not Written. It's Audited. Write discipline always loosens. Three weeks in, the log claims things are done that the state file still shows in flight — and nobody notices. The fix is not stricter discipline; it's a weekly mechanical diff between what your log claims and what your state shows. Playbook
- 2026-09-30 Microsoft put 1,024 AI agents on one task. What did it cost? Microsoft's Agensh paper scales one task from 1 agent to 1,024 with no orchestrator. The gains are real but sharply sublinear — and the token bill appears nowhere in the paper. Industry
- 2026-09-29 Your API key is probably already on GitHub Searching GitHub for OPENAI_API_KEY returns thousands of committed .env files — many still live. Five boring fixes that actually protect your keys. Guide
- 2026-09-29 It Remembers Because You Close Memory systems don't fail at install — they fail at the end of each session, one skipped write-back at a time. The fix is a two-minute closeout ritual, and it now ships as a paste-in prompt. Playbook
- 2026-09-28 The Authors Guild filings are an API procurement problem Unsealed briefs in Authors Guild v. OpenAI quote executives admitting they trained on pirated books. Two questions every API buyer should now ask their provider — provenance and retention. Industry
- 2026-09-25 Your AGENTS.md Is Rotting (Here's the Fix) AGENTS.md rots by default: stale decisions, invented paths, confident lies. Structure won't fix it — a maintenance layer will. Here's ours, in production. Playbook
- 2026-09-21 AI remembers, you find: fixing AI memory rot AI memory rots mid-project: settled decisions get relitigated, files get invented. Wallaby-agent-rules v2 moves memory to plain, dated, sourced files. Playbook
- 2026-09-14 How to compare LLM API costs: measure cost per accepted output LLM API cost comparison goes past price per million tokens: retries and human fixes are the bigger bill. Measure cost per accepted output in one afternoon. Playbook
- 2026-09-13 AGENTS.md as a token budget: the forbidden list AGENTS.md is a Linux-Foundation standard read by 30+ coding agents. Used well it's a token budget — our full production playbook, forbidden list included. Playbook