Mem0 reports 94.4% on LongMemEval. Zep, Letta, and a half-dozen newer memory services publish their own green numbers. And yet open any developer forum and the same complaints keep coming back: my agent forgot what we decided last week; it confidently executed a plan we abandoned; it remembered something I never approved.
Both things are true at once, and the reason is structural: AI memory is three problems, and the industry benchmarks only measure the easiest one.
The three problems, up front:
- Storage and retrieval: can you find the fact later? Largely solved, and largely commoditized.
- Freshness and consistency: is the fact still true? Does it contradict a newer fact? Half-solved at best, and nobody sells a fix.
- Governance: who decided this was worth remembering, who owns it, and what happens when it changes? Unsolved, and nobody even sells a claim.
If your agent's memory feels broken, it is almost certainly dying in layer two or three, while the vendor's scorecard sits in layer one, glowing green.
Disclosure: Wallaby Token operates an OpenAI-compatible inference API, and our own agents run on plain-file memory (the open-source wallaby-agent-rules repo we publish, MIT). We are a participant in this market, not a referee, so we'll show our receipts and link our sources, and you can check every number here.
The benchmark mirage
LongMemEval, LoCoMo, and their siblings measure a specific skill: given a haystack of past conversations, can the system retrieve the right needle? That is a real skill, and the teams posting these scores are doing real engineering. We are not going to pretend otherwise.
But watch what happens when a score gets sliced. The maintainer of ai-memory, an open-source agent-memory project we respect for its unusual honesty, publishes not just the headline number but the breakdown: hit@5 of 0.823 overall, and in the fine print, 0.570 on multi-session recall and 0.598 on temporal reasoning (source). The scenarios production users actually live in, many sessions and facts that change over time, are exactly the slices where everyone is weakest. The easy slices carry the average.
There is a second, quieter problem: retrieval benchmarks are scored on whether the system found the fact. They never ask whether the fact was still true. A memory system that perfectly retrieves a stale decision scores full marks while actively harming you. No mainstream benchmark deducts for that.
We don't say this from a distance. We run our own business on file-based memory, and our internal failure registry — every entry dated, sourced, and written when the failure happened — holds 45 logged memory failures as of this writing. Almost none of them were retrieval failures. Here is what actually broke.
Problem two: memory rots
Our failure log reads nothing like a benchmark test set. A sample of the categories:
- Stale but confident. A file said a decision was current; the decision had been reversed days earlier. The agent executed the old plan with full confidence, because nothing in the file said it had stopped being true. Without a date and a source on the line, "outdated" is unprovable.
- Claimed done, never done. Entries marked complete that the logs could not back up. This one is so common that we now have a mechanical scan for it, shipped in the public repo as
templates/reconcile.py, which diffs the work log against the memory files and flags entries that claim completion without evidence. - Silent drift. Two files that were supposed to say the same thing slowly diverging as edits landed in one and not the other. Nobody noticed until an audit diffed them.
- The truncation wall. An entry file grew past 32 KiB and the agent host started silently reading a shortened version. Nothing crashed; the system just quietly knew less. We only caught it because a health check watches file size.
Every one of these passed retrieval.
Notice the shape of these failures: they are all write-side and maintenance-side. The retrieval layer was fine. The memory was findable. It was just wrong, stale, contradictory, or truncated. This is what "memory is hard" actually means in production: less about finding things, more about keeping what you found true.
The fix is unglamorous: dated lines with sources, git history as provenance, a two-minute closeout ritual at the end of every session, and scheduled mechanical audits that exit non-zero when they find rot. We've written up the closeout ritual and the audit layer in earlier posts, and the whole setup is in the repo. None of it uses an LLM, because a rule doesn't get tired on day forty.
Problem three: nobody is in charge
The third layer is the one nobody ships an answer to at all.
Every memory system has to decide, at write time, what is worth keeping. In the hosted products, that decision is made by an extraction model you cannot inspect: an LLM reading your conversations and choosing for you, non-deterministically. When we benchmarked one self-hosted engine on our own documents, the same document distilled into 24 memory entries in one run and 57 in another. Extraction is a dice roll. In file-based setups like ours, the decision is made by whoever wrote the line, but then the questions multiply. Who approved promoting this to the permanent file? What happens when two sessions write at once? When a fact changes, does the old belief get invalidated, or does it just sit there looking exactly as confident as the fresh one?
Benchmarks don't touch any of this. Product pages don't either. And yet this is the layer where an organization's memory becomes either an asset or a liability, because a memory nobody is accountable for is just a rumor with a timestamp.
Why the industry is stuck on layer one
If layers two and three are where production pain lives, why does everyone compete on layer one? We see three structural reasons, and notably, none of them is "the big labs are slow" or "the startups are bad."
Benchmarks are legible; maintenance is not. A retrieval score fits in a tweet and a pricing-page comparison table. "We catch claims of completion that lack evidence" does not. Marketing follows what can be screenshotted, so engineering follows marketing.
The hosted model's incentives point the other way. One major cloud vendor's memory service spells it out on its public billing page: memory storage and memory operations are billed at fractions of a cent, and the meter that matters is the model calls (billing doc). Memory, in that business model, is a hook; the profit center is inference. A hook is not supposed to become the product, and a vendor whose business is selling tokens has no structural reason to help you spend fewer of them re-feeding context.
Scale pushes toward unaccountable design. Serving millions of end users forces stateless, automatic, nobody-in-the-loop extraction. The features that make memory trustworthy, a human gate on what becomes permanent and explicit invalidation when facts change, are exactly the features that don't scale to a managed multi-tenant API. This is a design-point conflict, not a talent gap. The hosted vendors made a rational choice; it just isn't a choice that produces accountable memory.
What we run instead
Our answer is the boring one, and it's been public the whole time: plain files the customer owns, dated and sourced lines, git as the audit trail, a closeout ritual, and mechanical audits on a schedule. It charges discipline instead of money: a dated line when you write something down, two minutes of review when you end the day. We laid out the full trade-off in Three Routes, Three Bills.
What we'd love to see from the field: a benchmark for staleness, where you give systems a fact, change the fact, and check which one comes back. A benchmark for contradiction, where two conflicting claims go in and we see if anything notices. A benchmark for accountability: can the system tell you why a memory exists, and who approved it? The first team to publish honest, sliced numbers on those axes will teach the industry more than another decimal point of recall.
Until then: when a memory vendor shows you a green scorecard, ask which layer it measures. Then ask your agent what you decided last week, and watch what happens.
The files, prompts, and audit scripts are MIT-licensed at wallaby-agent-rules. If you run your own failure registry, we'd genuinely like to hear what broke.