Two weeks after installing a memory system, most projects are back to normal — which is the problem. The files are all still there: the entry file, the long-term memory, the index. But the entries describe week one, the index points at files that moved on day nine, and the agent is once again asking questions you answered on day three. Nothing broke. You just stopped closing.
Installing memory is a thirty-second event. Keeping it true is a daily habit — and the habit is the part nobody ships. Our last post in this series argued that structure solves loading, not freshness, and that the fix is a maintenance layer. This post is the smallest piece of that layer, the one everything else depends on: what you do in the last two minutes of a work session. We call it the closeout, and we run it daily — on the same project our agents work on: code, deploys, monitoring, this blog.
Disclosure: Wallaby Token operates an OpenAI-compatible inference API. The memory system described here is open source, MIT licensed, and is the exact setup our own agents run on.
Where the rot actually happens
A work session produces residue: decisions made, facts discovered, files created, wrong turns taken. When the session ends, all of it goes one of three ways:
- It gets written back to the right file, dated and sourced.
- It gets left in the chat log, where it gets compacted away a session or two later.
- It gets dumped into memory wholesale — every scratch note, every abandoned approach — until the pile grows so high the agent can't tell your iron rules from your scratch notes.
Option one is the only good outcome, and it never happens by default. The agent's job was to finish the task; tidying up its own trail is nobody's job unless you make it somebody's. So the default outcome is some mix of two and three: you lose what mattered and keep what didn't. Fifty messages later, the memory rots — not because the design was wrong, but because sessions ended without a closeout.
The four questions
Our closeout is four questions, applied to everything the session produced. That's the whole mechanism:
- Did this change a lasting fact? → Write it to long-term memory (
MEMORY.md), one dated line with a source. Iron rules and key decisions live here; nothing here ages out. - Is it evidence for a future judgment call? → Append it to the dated log (
LOG.md). Append-only, never rewritten — a record of what was decided, and why, is exactly what you need when a decision gets revisited. - Is it just "this got done"? → One dated line in current state (
NOW.md), with a pointer to the details. Nobody re-reads the details; everybody scans the line. - Is it process scratch? → Delete it. This is the counterintuitive one, and the most important: intermediate notes, abandoned approaches, scratch work — gone. Git history already holds the trace if you ever need it. Process scratch is the largest source of context garbage we know of, and the instinct to keep it "just in case" is how memory files become junk drawers.
Two minutes, once, at the end of the session. The discipline is the point — none of the four questions is clever.
The fifth step: did the plan survive?
One addition we borrowed from spec-driven development and kept: before closing, compare what you set out to do with what actually happened — one line. When the two drift apart, that drift is itself a fact worth recording, because next session's plan will be built on what you think happened. Most memory systems record outcomes; recording the gap between intent and outcome is what keeps the outcomes honest.
Make it a word, not a workflow
Here is the failure mode we hit ourselves: the ritual existed, was documented, and was skipped anyway — because "run the closeout procedure" is a workflow, and workflows need someone to remember them. What fixed it was hanging the ritual on a word. We say "wrap up" (or "note this down," or "that's wrong — log it") and the agent executes the whole protocol: update the log, refresh current state, sort every conclusion into its file, run the drift check.
The lesson generalizes: the system should adapt to your language, not the other way around. The right trigger words are the ones you already say when you're done for the day — yours, not ours. In v3 the entry file carries a Ritual words section that defines exactly this mapping, and the interview path asks which words you actually use.
The mechanical backstop
Rituals have bad weeks, so the ritual gets a backstop that can't. The health-check script in the repo runs on a schedule and flags drift mechanically. In v3 it learned four new scans:
- Current-state entries that haven't moved in 30 days — stale by definition, time to archive.
- Index drift in both directions: files that exist but are registered nowhere, and registrations pointing at files that no longer exist.
- Entry files approaching the size wall — several agent tools truncate at 32KiB, and they don't warn you; your instructions just stop being read. We check bytes, lines, and estimated tokens, because CJK text hits the same byte wall roughly three times earlier than English, and a line count never sees it coming.
- Long-term memory lines with no date — without one, "idle for 30 days" is unfalsifiable and nothing ever moves.
Zero dependencies, exit code 0 when clean, 1 when findings. The script is dumb, and that is its virtue: it doesn't call in sick.
What's in v3
All of the above is now in wallaby-agent-rules, MIT licensed:
- L3 — Closeout ritual, a fourth paste-in prompt in
PROMPT.md: the four questions, the drift check, and the file updates, run on demand. - Ritual words in the entry-file template, plus a seventh interview question so the words are yours.
- A deeper
health_check.pywith the four scans above. - Filled-in examples of
NOW.mdandINDEX.md, so you can see what "good" looks like after two weeks, not just on day one.
Upgrading keeps the three rules: add, never overwrite; your content is sacred; nothing moves without your confirmation. The one exception is health_check.py itself — it's scaffolding with none of your content in it, so the upgrade check will offer to swap it outright. Version markers move to v3; the L2 upgrade prompt in PROMPT.md detects v1 and v2 installs and proposes only the diff.
The honest limit
A two-minute ritual feels too small to matter. Skip it twice and it stops existing — then you're back to trusting the context window, which is where the rot started. Everything in this post is deliberately boring: four questions, one drift line, a word you already say, a script with no dependencies. Memory systems don't die from bad architecture. They die at close of business, in two-minute increments, and boring is the only thing that survives contact with a Friday afternoon.
What we run, and what we publish
The file layer, the health-check script, and now the closeout prompt are open source in wallaby-agent-rules — the same files and the same ritual we use daily. The design principles are in AI remembers, you find; the maintenance layer this ritual plugs into is in Your AGENTS.md is rotting; the token-budget view of the entry file is in AGENTS.md as a token budget.
We did not publish a product, because there isn't one: no service, no sync, no lock-in. What we published is the part that survives being given away — plain files and the habits that keep them true.
Your code, your business
Three commitments, verbatim from our privacy policy: No content logs. No training on your data. No usage reports built from your traffic. Usage lines record token counts, costs, and timing — never prompts, never completions. Your memory files never touch our systems at all — they live in your repository, which is the entire point.
Reliability you can verify
We operate a public status page so you can verify availability independently before troubleshooting your own setup. Our terms are written in plain language and publicly accessible. Wallaby Token is a registered Australian company with an ABN on file, and we run our own development workloads through the same gateway we sell — the agents behind this post ran on it.
Get started
New here? The fastest path is the L0 prompt in PROMPT.md — paste two lines, watch your agent build the files and hand you the first project map. Already running the system? Paste the L2 — Upgrade check prompt, then make the L3 closeout the last thing your agent does today. A memory system is only as good as its last closeout.