Here is a failure mode nobody screenshots. Your project log says the checkout redesign shipped on Tuesday. Your current-state file still lists it as in flight, waiting on assets. Both files are "your AI's memory." One of them is lying, and nothing in the system is going to tell you which.
We found our own version of this the unglamorous way: not by noticing, but by writing a script that notices. It reads the dated log, reads the state file, and diffs what the one claims against what the other shows. The first run flagged entries we would have sworn were accurate. That is the entire argument of this post: write discipline always loosens, so memory needs an audit layer, not more discipline.
Disclosure: Wallaby Token operates an OpenAI-compatible inference API. The memory system described here is open source, MIT licensed, and is the exact setup our own agents run on.
The two lies memory files tell
Once we started looking, the drift came in exactly two shapes:
The contradiction. The log records a closure (shipped, fixed, merged, published) and the state file still carries the same topic as open work. Usually nobody updated the state file, because updating it was part of the closeout, and the closeout got skipped on a Friday. The log is append-only and honest; the state file is edited and forgetful. Left alone, the gap compounds until the state file is a museum of things you finished months ago.
The rumor. The log claims done, but points at nothing: no file, no ticket, no link. A closure you cannot verify is indistinguishable from one that never happened. This one is subtler, because the entry feels complete when you write it. Six weeks later, "fixed the retry bug" with no pointer is a rumor you told yourself.
Both are boring. That is what makes them dangerous: there is no dramatic failure, just a slow divergence between what your records claim and what is true.
Why mechanical, not motivational
The obvious fix is a sterner resolution to keep the files in sync. We tried that. It works for about two weeks. We know, because our audit script is how we found out it had stopped working.
So the rule we now run is: the human ritual writes, the machine audits. The closeout ritual (from our last post) is still the entry point; it sorts each session's output into memory, log, state, or the bin. But rituals have bad weeks, so a scheduled script checks the results. And it checks the one thing the ritual cannot check about itself: whether the story the log tells and the story the state file tells are still the same story.
The script is deliberately dumb. It does keyword matching, not semantic judgment, and it reports candidates rather than verdicts; a human decides which file is stale. Dumb is the feature. A smarter audit would need configuration, tuning, and eventually its own maintenance ritual. A dumb one just runs.
What's in v3.1
The audit ships today in wallaby-agent-rules as templates/reconcile.py: zero dependencies, same exit-code contract as the health check (0 clean, 1 findings), two scans:
- Closure contradictions: log entries that claim something is done while the same topic still sits in the state file's "in flight" section.
- Evidence-free closures: "done" entries with no path, ticket, link, or backtick reference to verify against.
New installs get it by default: the L0 and L1 prompts in PROMPT.md now build six files instead of five. Existing installs need exactly one step: copy templates/reconcile.py into your project as scripts/reconcile.py and run it weekly next to the health check. The L2 upgrade check detects v3.1 by the file's presence, and like health_check.py, the script is scaffolding: it holds none of your content, so upgrades can offer to swap it outright.
The honest limit
A keyword diff will flag some false positives: a closure whose topic overlaps unrelated open work, a pointer written in a format the script doesn't recognize. We accept that trade knowingly. The audit's job is to make drift visible, not to be right about it. A candidate you dismiss in three seconds costs three seconds. A contradiction nobody surfaces costs you the week you spent trusting a file that had quietly become fiction.
The harder limit is the one from the last post, unchanged: none of this survives if you never run it. Hang the audit on the same schedule as the health check, and let the machine be the part of the system that never gets tired.
What we run, and what we publish
The reconcile audit is the third layer of the open stack: the file conventions and health check from v2, the closeout ritual from v3, and now the consistency audit, all in wallaby-agent-rules, MIT licensed, all running daily on our own project. The design principles are in AI remembers, you find; the maintenance layer is in Your AGENTS.md is rotting.
Our internal version of this script checks more (size budgets, decision ledgers, registration drift), because our workspace grew those needs. What we publish is the part every project needs on day one: the diff between what you said you did and what your files say is true.
Your code, your business
Three commitments, verbatim from our privacy policy: No content logs. No training on your data. No usage reports built from your traffic. Usage lines record token counts, costs, and timing — never prompts, never completions. Your memory files never touch our systems at all; they live in your repository, which is the entire point.
Reliability you can verify
We operate a public status page so you can verify availability independently before troubleshooting your own setup. Our terms are written in plain language and publicly accessible. Wallaby Token is a registered Australian company with an ABN on file, and we run our own development workloads through the same gateway we sell — the agents behind this post ran on it.
Get started
New here? The fastest path is the L0 prompt in PROMPT.md: paste two lines, and the audit ships with the install. Already running the system? Copy templates/reconcile.py into scripts/ and run it once today. If it comes back clean, congratulations; your log and your state file agree. If it doesn't, you just found out what the audit is for.