Here is the only rule in this playbook: before you build any memory feature, ask which of three mechanisms it strengthens: consolidation, pruning, or myelination. If the answer is none, don't build it. A longer context window strengthens none of the three. Neither does another retrieval layer. The three mechanisms below are the entire difference between an agent that accumulates files and an agent that accumulates judgment.
Each part below is something we run in production on our own project, with the failure mode it survived. The full stack is open source, MIT licensed, and linked at the end.
Part 1: Consolidation — a replay job, not a bigger store. A scheduled job reads the day's append-only work log, scores every entry against a rubric, and writes only what passes to the long-term store. Details die; gist survives.
Part 2: Pruning — deletion as a first-class feature. Finished work removes its open tracking cards, removed cards are never resurrected, and untouched items decay on a timer.
Part 3: Myelination — procedures that get used get fast. Repeated checklists harden into executable reflexes with trigger words, so the paths you walk most cost the least context.
Part 1: Consolidation — the replay job
Biological memory does not store the day; it sleeps on it. During sleep, the hippocampus replays the day's activity and the neocortex keeps the compressed result. The engineering translation is a scheduled job, not a database:
- Append-only log. Every work session appends timestamped entries to one dated log file. Nothing is edited in place, ever. You cannot audit what you rewrite.
- Scheduled replay. A cron job (ours runs four times a day) reads the last three days of entries and scores each one against a rubric: did it produce a usable artifact, change a decision, or close a loop? High scores write straight to the project board. Borderline scores go to a stronger model for a second opinion. Low scores are judged and dropped. Not archived; dropped. Our last full run: 355 raw entries in, 37 kept.
- Dedup before write. New candidates are compared against existing cards by title similarity, so the same event cannot board twice.
The failure mode to design around: dedup by similarity misses paraphrases. Our threshold let "writing the tutorial" and "shipped the tutorial" coexist as separate cards (same topic, different words), and the stale twin survived for days. If your rubric is the only thing between reality and your long-term memory, a human still needs to glance at the result. We treat that glance as a feature: consolidation should be auditable, not invisible.
Part 2: Pruning — forgetting on purpose
A child's brain loses billions of synapses on the way to adulthood and comes out stronger. The design principle: fewer connections, better ones. Three mechanisms make deletion systematic instead of accidental:
- Closed-loop cleanup. When a finished card appears in the done column, any open card covering the same topic is removed. Completion should kill its own tracking artifacts. Otherwise your board becomes an attic of finished work nobody threw out.
- The freeze rule. Anything a human removes from the board is recorded as deliberately forgotten and is never automatically re-added. This is the difference between pruning and amnesia: pruned items stay pruned.
- Decay with visibility. Cards untouched for thirty days get demoted out of the active view; cards on the board carry an age badge so staleness is visible at a glance. Memory that does not decay is hoarding.
The failure mode here is the opposite one: pruning that is too aggressive deletes things that were true but quiet. We audit the dropped pile on a schedule, precisely because it is a pile of deletions. Prune, then count what you pruned.
Part 3: Myelination — procedures into reflexes
The third mechanism is about speed, not storage. Neural pathways that fire repeatedly get myelinated and conduct faster; the procedural version is that a checklist you run every week should stop costing you reasoning.
The implementation is a lifecycle, not a document: register, verify, harden, retire. A new procedure starts as a written checklist. Once it survives a few real runs, it hardens into an executable script or a skill with a trigger word, so invoking it costs one line of context instead of fifty. Procedures that stop being used get retired out of the index, which keeps the reflex layer itself pruned.
The tell that you need this: your agent keeps re-deriving a process it has already run correctly. Every re-derivation is context spent re-learning something that should have been a reflex.
The board was never a kanban
One piece of infrastructure ties all three together, and we did not understand what it was until it broke. Our project board is not a task tracker; it is an external memory organ. The backlog is prospective memory (remembering to do things, the kind human memory is worst at). The doing lane is working memory, capacity four chunks plus or minus one, which is why it breaks the moment you overload it. The review lane parks unfinished business so it stops nagging. Bluma Zeigarnik documented that nag in the 1920s. The done column is the episodic archive.
Once you see the board as memory, the design questions change. "Should this be on the board?" becomes "is this worth encoding?" "Should we remove this card?" becomes "is this ready to be forgotten?" The narrative version of this story, including the stale card that started it, is in Your Agent Needs Sleep, Not a Bigger Context.
What we run, and what we publish
Everything above runs on our own project daily, and the publishable core is in wallaby-agent-rules, MIT licensed: the file conventions, the closeout ritual that feeds the log, and the audit layer that checks the whole machine. The audit side is in Memory Is Not Written. It's Audited.; the write-back discipline is in It Remembers Because You Close.
Disclosure: Wallaby Token operates an OpenAI-compatible inference API. The memory system described here is the exact setup our own agents run on.
Your code, your business
Three commitments, verbatim from our privacy policy: No content logs. No training on your data. No usage reports built from your traffic. Usage lines record token counts, costs, and timing — never prompts, never completions. Your memory files never touch our systems at all; they live in your repository, which is the entire point.
Reliability you can verify
We operate a public status page so you can verify availability independently before troubleshooting your own setup. Our terms are written in plain language and publicly accessible. Wallaby Token is a registered Australian company with an ABN on file, and we run our own development workloads through the same gateway we sell — the agents behind this post ran on it.
Get started
New here? The fastest path is the L0 prompt in PROMPT.md: paste two lines, and the log-and-closeout backbone ships with the install. Then take the rule from the top of this post and apply it to the next memory feature anyone proposes: consolidation, pruning, or myelination. Name one, or don't build it.