Your Agent Needs Sleep, Not a Bigger Context

This afternoon my partner glanced at our project board, pointed at one card, and said: "That one is obviously wrong." The card tracked a tutorial our agent was supposedly still writing. The tutorial had shipped six days earlier — published, indexed, live.

Nothing had crashed. No alert fired. The system did exactly what we built it to do: remember everything. And that was the bug. It remembered the tutorial as in-flight because nobody had explicitly told it otherwise, and a system that never forgets has no way to notice that reality has moved on.

We had spent months building our agent a better memory, and we had built the wrong thing. Not because the memory was inaccurate. Because it was complete.

Everyone is building a bigger notebook

The entire industry's answer to agent memory is the same: store more, retrieve better. Longer context windows. Bigger vector databases. Another retrieval layer on top of the last one. The implicit theory is that memory failure means recall failure, so the fix is always more storage with better search.

That is the notebook theory of memory, and it is wrong in a specific way: a notebook never gets smarter. Everything you write in it sits there forever with equal weight, competing for your attention every time you open it. A perfect notebook is a perfect hoarder.

Ask why humans don't work this way and the whole design space changes.

Human memory's main job is forgetting

Your brain is not trying to remember your day. Tonight, while you sleep, it replays the day's experiences at high speed, keeps the gist, and lets the details die. That process has a name: systems consolidation, the transfer from hippocampus to neocortex. It behaves less like backup and more like editing. The price you paid for lunch is gone by Friday; the fact that the client hated the pricing page is not.

Then it deletes. Synaptic pruning removes weak connections outright. A child's brain loses billions of synapses on the way to adulthood, and comes out stronger for it. Fewer connections, better ones.

And for what survives, it builds fast paths. Repeated activation wraps neural fibers in myelin and signals travel up to a hundred times faster. Practice doesn't just make perfect; it makes myelin.

Consolidation, pruning, myelination. Notice what none of them do: none of them store more. Human memory is an accuracy-and-strength machine, and its raw material is deletion.

The system we had already built by accident

Then came the uncomfortable part. Once we laid the three mechanisms on the table, we found we had built pieces of all three. We just never named them.

Four times a day, a scheduled job reads the day's work log, scores every entry against a rubric, and decides what gets written to the project board and what gets dropped. The last full run read 355 raw entries and kept 37. The rest were judged, not lost. That is consolidation: the details die, the gist floats up. We run it on a timer, which makes it, literally, sleep replay.

Then there is the freeze rule: once a card is removed from the board, the system never re-adds it. Forgotten stays forgotten. A cleanup pass removes an open card when a finished card already covers the same ground. That is pruning.

And our SOPs started life as careful checklists and are now reflexes: the deployment sequence, the release gates. That is myelination. The path is the same every time, so it got fast.

Even the board itself turned out to be a memory organ wearing a kanban's clothes. The backlog is prospective memory, the remembering-to-remember kind, the one humans are worst at. The doing lane is working memory, which is why it breaks the moment it holds more than a handful of cards: human working memory holds four chunks, plus or minus one, and Nelson Cowan measured that decades ago. The review lane exists because of Bluma Zeigarnik, who noticed in the 1920s that waiters remember open orders in vivid detail and forget closed ones instantly. Unfinished business nags until it is parked somewhere trusted. And the done column is the episodic archive: what happened, when, who decided.

We thought we had built a place to sync progress. We had actually built a shared external memory, what psychologist Daniel Wegner called a transactive memory system: the thing couples and teams do when nobody remembers everything but everybody knows where the index is.

Where it still fails honestly

The board caught the stale card because a human glanced at it. The machine had not. Our cleanup pass matches cards by title similarity, and "writing the tutorial" versus "shipped the tutorial" scored below the threshold. The twin cards were never linked, so the stale one survived until a pair of eyes walked by. Pruning, it turns out, has a last inch that is still human.

The scoring thresholds are argued, not derived. We raised the bar when too much noise got through and lowered it when real work got dropped, which is a fancy way of saying the system's sense of what matters is calibrated by disagreement. And consolidation is lossy by design. Every run throws away entries that were true, just not durable. We audit the deletion pile precisely because it is a deletion pile.

We are telling you this because the posts that only show the working parts are how everyone ended up believing memory was a storage problem.

The rule we kept

Out of the wreckage we kept one sentence, and it now gates every memory feature we consider:

Before you build it, ask which of the three it strengthens: consolidation, pruning, or myelination. If the answer is none, don't build it.

A bigger context window strengthens none of the three. Neither does another retrieval layer. A nightly job that reads the day's work and throws most of it away strengthens all three at once.

Your agent does not need to remember more. It needs to sleep on it.

What we run, and what we publish

The memory stack behind this post is open source and MIT licensed: the file conventions, the closeout ritual, and the audit layer are all in wallaby-agent-rules, running daily on our own project. The board in this post is the live one. The audit side of the system is in Memory Is Not Written. It's Audited.; the write-back discipline is in It Remembers Because You Close.

Disclosure: Wallaby Token operates an OpenAI-compatible inference API. The memory system described here is the exact setup our own agents run on.

Your code, your business

Three commitments, verbatim from our privacy policy: No content logs. No training on your data. No usage reports built from your traffic. Usage lines record token counts, costs, and timing — never prompts, never completions. Your memory files never touch our systems at all; they live in your repository, which is the entire point.

Reliability you can verify

We operate a public status page so you can verify availability independently before troubleshooting your own setup. Our terms are written in plain language and publicly accessible. Wallaby Token is a registered Australian company with an ABN on file, and we run our own development workloads through the same gateway we sell — the agents behind this post ran on it.

Get started

New here? The fastest path is the L0 prompt in PROMPT.md: paste two lines, and the consolidation layer ships with the install. Then do the uncomfortable part: find one thing your agent remembered that stopped being true, and delete it on purpose.