AI Memory: Three Routes, Three Bills

This morning we pointed a popular open-source memory engine at our own project memory — four real files, fifteen questions. It answered twelve, at an average of 19 milliseconds, including all four questions where grep scored zero. By lunch we had decided not to use it.

That "no" is what this post is about. There are three routes to giving an AI agent memory, and after spending the day measuring one of them against our own setup, we're convinced the routes differ less in technology than in what they bill you. Pick the route whose bill you're actually willing to pay.

The three routes, up front:

  • Route A — hosted memory API. Send conversations to a memory service, query it later. Bills you money and trust.
  • Route B — self-hosted memory engine. Run the embedding-and-extraction pipeline on your own box. Bills you engineering hours and supply-chain audit.
  • Route C — plain files plus conventions. Markdown files, an index, and rules about what gets written where. Bills you discipline, every single day. (This is the route we run — the files and scripts are open source.)

None of them is free. Two of them don't look like they cost anything, which is worse.

Disclosure: Wallaby Token operates an OpenAI-compatible inference API. Our own agents run on Route C — the open-source, MIT-licensed file system we've written about in this series — and Route A and B products could, in principle, be built on top of an API like ours. We'll try to be fair anyway, and we'll show our work.

Route A: the hosted memory API

The pitch is genuinely good: one API call to store, one to recall, and the service handles extraction, embedding, ranking, and user profiles. No vector database to provision, no chunking strategy to argue about. Supermemory's cloud and Mem0's platform are the two names most people run into first.

The bill has two lines.

Money. Predictable, invoice-shaped, and honestly the least interesting line on the bill.

Trust. This is the line that matters, and it's easy to miss because nothing breaks when you sign it. Your agents' memories — which include your decisions, your customers, your half-formed plans — now live inside someone else's system. What gets extracted from a conversation, what gets ranked into context, what gets quietly dropped: all of it happens inside a pipeline you can't inspect, running on proprietary models you can't audit. Their docs are candid about this — the platform tier is exactly where the proprietary extraction models live.

For a todo-list app, none of this matters. For a system whose memory is the business — your ledger of what was promised to whom, and why — you're outsourcing the part you most need to verify.

Route B: the self-hosted engine

Route B keeps the machinery but brings it home: same pipeline, your machine, your keys, your choice of model. Supermemory's local server is the cleanest example right now — one binary, encrypted local storage, local embeddings by default, bring your own LLM over any OpenAI-compatible endpoint. So we did what Route B asks of you: we ran it.

Here's the honest scoreboard, because anyone weighing Route B deserves real numbers. We ran two rounds on our own documents: R1 with the stock English embedding model against a single file, R2 with a multilingual embedding model against four real project files — a project overview, a README, a slice of a decision log, a style baseline. Our grep-and-index setup answered the same questions as the baseline.

Question type grep + index Engine R1 (English embeddings, 1 file) Engine R2 (multilingual, 4 files)
Precise facts 8/8 4/8 4/6
Fuzzy semantics, different wording 0/3 — zero word overlap 2/3 4/4
Update conflicts ("which value is current?") — — 2/2
Multi-hop combinations — — 2/3
Total — 6/11 12/15 (80%)

Average query latency: 16 ms in R1, 19 ms in R2. (The two rounds used different question sets as the corpus grew — read R1 as a smoke test and R2 as the real measurement.) What we did not test: anything beyond a one-day, four-file, single-operator pilot — no multi-week drift, no concurrent writers, no scale, no adversarial junk in the corpus.

  • The multilingual swap was the single decisive change. Same pipeline, same documents, and fuzzy recall went from 2/3 to 4/4 purely from the embedding model. Evaluate a memory engine with an English-only embedding model on non-English documents and you're measuring the wrong component.
  • The actual questions, translated from our Chinese test set. "How do we bill customers?" and "what happens if a supplier stops delivering?" both hit — with zero shared keywords between question and source text. "What memory tools have we been evaluating lately?" missed, because the log entry holding the answer never used those words. And a company registration number sitting in a dense table row missed even when the query contained the number itself. Chunking is still lossy; precise identifiers belong in an index, not an embedding. Hence the grep column's 8/8 — and its 0/3 on differently-worded questions, where the engine went perfect. The engine and the index don't compete. They cover each other's blind spots.
  • Update conflicts went 2/2 — but watch what "wrong" would look like. Asked which of two values is current, the engine picked the newer chunk both times. Encouraging. But a chunk store has no visible notion of "superseded": when it gets this wrong, the stale value looks exactly as confident as the fresh one. Our dated lines answer "which is current?" by construction — the old value stays in the file, marked replaced, with the date it stopped being true.
  • Extraction is non-deterministic. Its LLM pass — running on our model, through our endpoint, not theirs — distilled the same overview doc into 24 atomic memory entries in R1 and 57 in R2. The auto-generated profile summaries looked sensible on inspection. And two of our four documents made the memory-agent step fail outright with request is invalid — an endpoint-compatibility quirk we didn't chase, since chunk search was unaffected.

One more thing the scoreboard doesn't show: the write path. Search quality is what everyone benchmarks, but a memory system's character is set at write time — what it decides is worth keeping. The engine's extraction pass is an LLM reading your documents and choosing for you; run it twice and it chooses differently. It has no way to know that one file is a scratch note and another is the record you'll need in a dispute. Our file setup makes that choice deliberately: permanent facts and iron rules live in one file, current state in another, scratch work gets deleted, and nothing is written without a dated line and a source. Slower to write. Impossible to misunderstand.

That's the good news, and it's real. Now the bill, which is also real:

  1. Engineering hours. The install was clean — we read the 333-line install script first, and it checks out. But then: the server died with ECONNRESET until we realized our proxy environment variables were hijacking calls to our own gateway; the 282 MB multilingual model stalled mid-download and had to be hand-placed from a mirror; a search parameter silently returned nothing until we noticed the API wants containerTags (array), not containerTag. A normal day of integration work — roughly one engineer-day, all in.
  2. Evaluation burden. The engine's cloud benchmarks don't transfer to a self-hosted setup, because the cloud runs proprietary extraction models and you don't. So you must build your own test set, from your own documents, and re-run it every time you change a model. That fifteen-question set cost us more thought than the install did. If you want to see what that burden looks like when someone carries it honestly, Nowledge's write-up of their memory-pipeline experiments is the best public account we've found — ablations, per-task costs, an earlier gain they explicitly retract when it fails to reproduce, and a shadow-mode rollout that records the new judge's answers while the old one stays in charge. (Full circle: their chat-model ablations ran on Kimi k3.)
  3. Supply-chain audit. The self-hosted server ships as a prebuilt binary. The repository is MIT licensed — we read the LICENSE file, and MIT is as permissive as it gets. The server's own source, though, isn't in the repo; issue #1299 asks where it lives, and the question is still open. None of which makes the software bad — it made our day measurably better. But our house rule for anything that gets read access to project memory is stricter than any license: if we can't see what it's built from, it doesn't get the keys. That bar is about us, not about them.

Route B's bill, in short: the software is free, and the confidence is not.

Route C: plain files, plus discipline

Route C is what our agents actually run on: an entry file, a long-term memory file, a current-state file, an index, and conventions about what gets written where and when. No embeddings, no pipeline, no service. We've published the whole thing — wallaby-agent-rules, MIT — and the design logic fills three earlier posts: structure solves finding, not freshness, the maintenance layer, and the closeout ritual.

Let's kill a misconception first, because it comes up a lot: Route C is not "the AI just remembers if you ask nicely," and it is not willpower. Willpower is the entry point, sure — but the system stands on two mechanisms that don't depend on anyone's mood:

  • Post-verification. A scheduled health-check script scans the memory files mechanically: entries that haven't moved in 30 days, index drift in both directions, entry files approaching the 32 KiB truncation wall that several agent tools hit silently, undated lines that make "stale" unfalsifiable. It exits non-zero when it finds something. It has no bad weeks.
  • Audit. Every memory line carries a date and a source; the whole thing lives in git, so "when did we start believing this, and who said so" is a git log away. You cannot diff an embedding.

The bill, honestly stated: discipline, daily. A two-minute closeout ritual at the end of every session, or the files rot. A weekly reconciliation pass, or the index drifts from reality. Route C doesn't bill you money and it doesn't bill you trust; it bills you two minutes at the end of every workday, forever, and it collects relentlessly. Miss it twice and the ritual is dead.

The bill nobody publishes

Here's what the day of benchmarking actually taught us, and it's the part we'd write even if nobody read it.

We keep an incident log for our own memory system — every time the agent forgot something it should have known, with the cause attached. Reading that log against the benchmark results produced an uncomfortable asymmetry: almost none of our real-world memory failures were retrieval failures. The retrieval layer — humble grep over humble files — nearly always had the answer. The failures were that the fact was never written down, or the file was never loaded, or the session ended without a closeout and the decision evaporated with the context window.

The industry publishes recall benchmarks obsessively. Nobody publishes a follow-through benchmark. And yet the evidence of our own operations says follow-through is where memory systems actually die — not with a retrieval miss, but in the two skipped minutes at the end of a workday.

Yifei Fang made the adjacent point on NVIDIA's engineering blog this week, writing about AI-generated code: "AI increases the rate at which candidate implementations can be produced. Architecture and validation determine whether that increased output becomes reliable software" — and the advice is to treat automated validation as the production constraint. Generation is cheap; verification is the bill. Memory is the same shape. Storage is cheap, recall is cheap, extraction is cheap. Knowing your memory is true — that's the production constraint, and every route bills it differently.

So which route?

Side by side, the three bills:

A · hosted API B · self-hosted engine C · files + discipline
You pay in money + trust engineering + audit daily discipline
Memories leave your perimeter yes no no
You can inspect ranking & extraction no partly — if the source is public fully — it's text
Fuzzy cross-file recall, out of the box yes yes no
Who decides what's worth keeping their pipeline an LLM pass you prompt but don't control you, at write time
What a wrong answer looks like confident, no age confident, no age dated and sourced — or absent
"When did we start believing this?" their dashboard you build it git log
Dies when their roadmap moves your eval set rots you skip the ritual

Not a verdict — a routing table:

  • Pick Route A if you're shipping a product feature on a deadline, the memories aren't sensitive, and you have no appetite for owning an eval set. Pay the money-and-trust bill with eyes open.
  • Pick Route B if you need data residency and you have a spare engineer-day plus the discipline to maintain your own test set. Audit the supply chain like you'd audit a dependency that can read your mail, because it can.
  • Pick Route C if the content is operational truth — decisions, promises, ledgers, the things you'd need to defend in a dispute — where auditability beats fuzzy recall. Or do what we do: C as the base layer that owns truth, with B-shaped retrieval as an optional add-on for the one case grep can't see.

And us? We said no to the binary, and yes to the hole it exposed. The engine's 4/4 on fuzzy cross-file questions proves the gap in our setup is real, so we've kept our measurements as the acceptance bar: if "I remember the gist but not the keyword" bites us twice more, we build the thirty-line embedding index ourselves — and it has to beat 12/15 at 19 ms before it earns a place next to the files. Until then, the index stays authoritative.

The memory route you want is the one whose bill you can see. The dangerous ones are the ones that look free.

What we run, and what we publish

The file layer, the health-check script, and the closeout prompt are open source in wallaby-agent-rules — the same files our agents ran while writing this post. The benchmark numbers above come from a one-day pilot on our own documents; treat them as one team's measurement, not a product review. If you're pricing the hosted options, we keep a current comparison of all four official pricing pages.

Your code, your business

Three commitments, verbatim from our privacy policy: No content logs. No training on your data. No usage reports built from your traffic. Usage lines record token counts, costs, and timing — never prompts, never completions. Your memory files never touch our systems at all — they live in your repository, which is the entire point.

Reliability you can verify

We operate a public status page so you can verify availability independently before troubleshooting your own setup. Our terms are written in plain language and publicly accessible. Wallaby Token is a registered Australian company with an ABN on file, and we run our own development workloads through the same gateway we sell — the agents behind this post ran on it.

Get started

New here? The fastest path is the L0 prompt in PROMPT.md — paste two lines, watch your agent build the files and hand you the first project map. Already running the system? Paste the L2 — Upgrade check prompt, then make the L3 closeout the last thing your agent does today. A memory system is only as good as its last closeout.