Every AI memory system starts trusting itself on day one. From the first session, the agent decides what is worth remembering, writes it down, and reads it back later with full confidence, before you've seen whether its judgment is any good.
We run file-based memory on our own business and log every failure: 47 entries so far (as of October 2026), each dated and sourced, with the full taxonomy in the open. The pattern that surprised us: the failures that hurt were almost never retrieval failures. They were bad writes — a wrong fact stored with total confidence, a decision recorded without its reason, a stale plan executed because nothing marked it as withdrawn. Retrieval gets benchmarked. Writing is where memory actually breaks.
So wallaby-agent-rules 1.0.0, the open-source memory setup we've been building in this series (structure, maintenance, audit), ships with the write side gated by default. We call it trust-building mode.
What it does
For the first two weeks after install, the memory system works exactly as before with one exception: new long-term memories are proposed, not written.
At the end of each session, the agent lists what it would have added to MEMORY.md, each entry marked pending, dated, and with a source, and waits. You reply "yes" to the ones that deserve permanence, and only those get promoted. Everything else expires quietly.
Then, two weeks in, the mode expires on its own, or whenever you say "trust mode off", and the system switches to direct writes. The probation period is over because by then you have something better than a default: two weeks of evidence about what your agent chooses to remember.
That distinction is the whole point. A permanent approval gate would mean you don't trust the system, full stop. A probation period means trust is calibrated on evidence: you watch the agent's judgment before you fund it with authority. New hires get probation. New memory writers should too.
What else is in 1.0.0
- A code-version memory module, mounted conditionally. The setup interview now asks whether your project involves code. If yes, the entry file gains a
Code & releasessection: the currently deployed version becomes a dated memory fact, updated the moment it changes, and every release, migration, or rollback gets logged with its commit reference. If no, the section is skipped entirely; a research project's memory shouldn't carry code discipline it will never use. - A full Chinese edition.
README.zh-CN.mdandPROMPT.zh-CN.md: the whole system in Chinese, with ritual words that are natively Chinese (收工 / 记一下 / 这不对), because the system is supposed to adapt to your language, not the other way around. - Semantic versioning, reset to 1.0.0. Earlier iterations, v1 through v4.1, are folded in. Patch releases are fixes, minors are features, and per our own release checklist, patches never get a Release of their own.
- A license change worth noticing. The repo is now MIT + Commons Clause: free for personal projects, research, and use inside your own organization; selling it needs a commercial license (COMMERCIAL.md). Your memory files live in your repository and remain yours, whichever tier you ever touch.
Try it
New install, about 30 seconds, defaults including trust-building mode on:
Fetch https://raw.githubusercontent.com/Dawncoral/wallaby-agent-rules/main/PROMPT.md
and follow its "L0 — One-click install" section in this project.
Already running an earlier version? Paste the L2 — Upgrade check prompt from the same file. It detects your version, proposes the additions item by item, and never overwrites anything you wrote; that rule hasn't changed since the first commit and won't.
And if you run your own memory setup with a different answer to the write-side problem, whether that's no gate, a permanent gate, or something smarter, we'd genuinely like to hear it. The repo is wallaby-agent-rules, and the failure registry that shaped this release is the reason it exists.
Your data, your business
Three commitments, verbatim from our privacy policy: No content logs. No training on your data. No usage reports built from your traffic. Usage lines record token counts, costs, and timing — never prompts, never completions.