Shopping for an AI memory layer in 2026? Here are the four hosted options you'll run into first, priced from their official pages today (October 1, 2026), plus two routes that don't appear on any pricing page: self-hosting an open-source engine, and plain files. We ran the self-hosted route on our own documents last week, so the numbers below include one set nobody else publishes: what that route actually costs in effort.
Disclosure: Wallaby Token operates an OpenAI-compatible inference API. Any of these products could, in principle, run their extraction on an API like ours, and our own agents run on the plain-files setup described at the end. Every price below links its source.
The four hosted APIs, priced — plus ours
| Free tier | First paid | Upper tiers | Meters | |
|---|---|---|---|---|
| Supermemory | $5 of credits monthly | Pro $19/mo ($20 credits) | Max $100, Scale $399 | Tokens ingested + queries (rates) |
| Mem0 | 10,000 adds / 1,000 retrievals monthly | Starter $19/mo | Pro $249 (graph memory starts here) | Add + retrieval requests (pricing) |
| Zep | 10,000 credits monthly | Flex $125/mo (50,000 credits) | Flex Plus $375 | Credits per ingested episode; retrieval unmetered (pricing) |
| Letta | 3 stateful agents, BYOK | Pro $20/mo (20 agents) | Teams $20/seat | Agents + LLM usage; self-hosting free (pricing) |
| Plain files (what we run) | Everything, forever | — | — | Nothing. Your files, your git repo (ours is open source) |
The catch: the four meter different things. Supermemory bills tokens in and queries out. Mem0 bills requests in both directions, and its 1,000-retrieval free ceiling arrives long before its 10,000-add one. Zep bills only what you send in; retrieval, storage and users are unmetered. Letta bills per active agent plus the underlying model spend. A workload that's cheap on one rate card can be expensive on another, so model your own retrieval volume before comparing headline prices.
Two details worth knowing: Mem0's graph memory starts at $249/month, and Zep's enterprise tier is the one with SOC 2 Type II and HIPAA. Supermemory's Scale tier ($399) is where managed self-hosting enters; its open-source local server is free today.
The line item no pricing page shows
A hosted memory API decides what your agent remembers (extraction, ranking, pruning) inside a pipeline you cannot inspect. When it silently drops a fact, the answer you get looks exactly as confident as a correct one. That's not an argument against buying; it's an argument for knowing what you bought. We wrote the long version in a companion piece.
The route no pricing page lists: self-hosting
"Just self-host the open-source engine" is the answer forums give. We did it: one binary, encrypted local storage, our own model over our own endpoint, fifteen questions against four real project files. The scoreboard: 12/15 answered correctly, at 19 milliseconds average per query, including 4/4 on fuzzy, differently-worded questions where our grep-and-index setup scores zero.
The bill came in a different currency. Roughly one engineer-day of integration: a proxy environment variable hijacked calls to our own gateway, a 282 MB model download stalled mid-way, an API parameter silently returned nothing until we noticed it wanted an array rather than a string. Then the ongoing costs: a test set you build and maintain yourself, and a supply-chain question we couldn't close, since the server's own source isn't in the repository and the issue asking where it lives is still open. Itemized in the full write-up.
Free software. Not a free lunch.
And the route with no vendor at all
The setup our own agents run: markdown files, an index, and conventions about what gets written where. No embeddings, no pipeline, no service. Permanent facts in one file, current state in another, every line dated and sourced, the whole thing in git. It costs nothing and trusts no one; it bills you two minutes of discipline at the end of every workday, forever. No fuzzy recall out of the box, which is exactly the gap our 4/4 measurement exposed. But for operational truth (decisions, promises, ledgers) auditability beats fuzzy recall, and you can't diff an embedding. Open source, MIT: wallaby-agent-rules.
Route yourself
- Supermemory if you want the most usable free prototype and a public metered rate card.
- Mem0 if you want the largest ecosystem and can live inside per-request tiers.
- Zep if retrieval volume is heavy (it's unmetered) and enterprise compliance is on the roadmap.
- Letta if you want a full agent runtime, BYOK, and a free self-hosting path.
- Self-hosted engine if data residency is non-negotiable and you have an engineer-day plus the patience to own an eval set.
- Plain files if the content is operational truth, or as the base layer that owns truth with an engine bolted on later.
Frequently asked questions
"Which AI memory API is cheapest?"
At entry level, Supermemory and Mem0 tie at $19/month, and Letta's Pro is $20. The real answer depends on what you meter: Zep's Flex ($125) stops charging for retrieval entirely, which beats a low headline price for read-heavy workloads. Run your own request volumes against each rate card; all four free tiers exist exactly for that.
"Can I self-host any of these?"
Three ways in. Supermemory's local server and Mem0's open-source framework are free to run on your own machine today. Letta is free to self-host with every feature. Managed self-hosting inside your cloud is an upper-tier or enterprise conversation at Supermemory (Scale, $399/month) and Zep (BYOC, custom), per their pricing pages as of October 2026.
"Is there a free option that's actually production-usable?"
For prototyping, all four hosted free tiers are real. For production with zero vendor, the two honest answers are a self-hosted engine (free software, real engineering) and plain files (free everything, real discipline). What neither gives you out of the box is fuzzy cross-file recall; that's the one thing worth paying for, whichever route you pick.
Where your memories live
One question cuts through every table above: who physically holds your agent's memories? On the hosted routes, the vendor does. On the file route, nobody does — they sit in your repository, which is the entire point. For what it's worth, our own API business runs on three commitments, verbatim from our privacy policy: No content logs. No training on your data. No usage reports built from your traffic. Whichever route you pick, ask your vendor the same questions.
Try the file route in two minutes
The L0 prompt in PROMPT.md: paste two lines into your agent, watch it build the files and hand you the first project map. The best memory system is the one whose answers you can check.