Layered Memory for Grok Build: Native Notes + Repo Files for Engineering Teams

If you've ever spent 10 minutes re-explaining your project's pytest config, commit message conventions, and architectural decisions to an AI coding agent — only for it to forget every single rule the next time you open a new session — you're not alone. And it's not just 10 minutes lost; it's inconsistent code, unlogged decisions, and silent drift across your team.

This is why we built a file-based memory system: we think memory is infrastructure for small engineering teams, not a feature you rent from one AI tool. It lives in your repo, travels with git, and works no matter which agent you use.

On September 16, Grok Build launched its own native memory feature, closing the cross-session forgetting gap at the personal layer. So we did what infrastructure people do: layered our repo-file system on top of it, and ran both configurations through the same five tasks to map where each layer's job ends. Same slug-generation project, two runs: one on native memory alone, one on native memory plus our three-layer repo files (AGENTS.md, layered memory files, an arbitration chain).

The two runs separated the layers cleanly:

  • Native memory alone held every convention across sessions, but on its own, none of the changes from four working sessions reached git
  • With the repo-file layer on top, 100% of the work was committed by end-of-day closeout, memory audit trail in the same commits

The wildest part? When we hard-reset the repo mid-test, the agent kept working off code that no longer existed. The repo forgot. The engine remembered.

28 minutes of session time, 12 screenshots, full git history preserved. Here's the complete breakdown.

The setup

The test project is slugtool, a small Python CLI that turns titles into URL slugs, with a unit test suite and zero dependencies. Five tasks, scripted verbatim in advance: set standing conventions (canonical check is python3 -m unittest -v, commit messages lowercase imperative, max 72 chars), make a design decision about non-ASCII titles and record the why, then a deliberate contradiction (the canonical check is now pytest -q, and rename the flag), wrap up for the day, and finally a next-day handoff in a fresh session asking for a --max-len option that follows the project's conventions.

Two runs. Run A used stock Grok Build 1.0.46 with its native memory enabled. Run B used the same binary plus two files from wallaby-agent-rules 1.0.1: an AGENTS.md entry file and a three-tier MEMORY.md (Active, Standby, Dormant). Same machine, same model, grok-4.7 on high, running on a grok.com subscription.

One setup finding worth knowing before anything else: native memory is off by default. You enable it with two lines in ~/.grok/config.toml, and the notes actually land in ~/.grok/memory-v2/, not the ~/.grok/memory/ path the docs mention. Each project gets its own workspace scope, keyed by directory name plus a hash, so our two runs were cleanly isolated from each other.

Grok Build 1.0.46 first launch in the Run A project

Run 1: the native layer alone

Session one went well. The agent added the --sep flag, validated it, wrote tests, all 15 passing, and then committed without being asked, with a message that followed the convention we had just stated. Native capture wrote two observations recording both standing rules. Zero configuration, automatic, genuinely convenient.

Session two made the design call: transliterate accents via NFKD rather than stripping non-ASCII. The decision landed in a topic file with the full reasoning, including edge cases for ß, æ, and CJK characters. Solid work. No commit, though. Three modified files stayed in the tree.

Session three was the contradiction. The agent obeyed the new check command instantly, with no arbitration and no confirmation. It renamed --sep to --separator as a breaking change, no alias. Pytest wasn't installed, so it ran pip install --user pytest, dropping pytest 8.4.2 into ~/Library/Python/3.9/bin, outside the project, on its own initiative. And the memory inbox kept the old convention notes right next to the new ones: eleven observations by the end of the run, "check is unittest" and "check is pytest" filed side by side. Which one wins at recall is opaque. xAI's own announcement states the rule honestly: "Instructions in the current conversation take precedence over anything in a note." Last word wins, every time.

Session four, "wrap up," took 28 seconds. It ran the check, gave a verbal summary, and wrote nothing new to memory. Three sessions of work sat uncommitted in the working tree.

Then day two happened.

Session five, the fresh-session handoff, is where native memory showed real strength: the agent used the new pytest -q convention rather than the old one, used --separator, and implemented --max-len correctly with word-boundary backtracking, 24 tests passing. Cross-session recall works. But it still didn't commit. The final git log for the entire run: one commit, from session one. Four sessions of changes, 0% traceability.

Run A session 1: unprompted commit, native capture writing standing conventions

Phantom memory: a design boundary, and the habit that covers it

We hit this one during setup, and it turned out to be the most instructive finding of the day. After a pilot session, we hard-reset the repo to drop the pilot's commit and start clean. The repository forgot. Grok's memory did not: three observations describing the discarded commit were still sitting in the workspace store. When the formal session one finished, six near-duplicate observations coexisted in the same inbox, half of them pointing at code that no longer existed anywhere in the repo.

Nothing flags this. There is no drift check comparing what the engine remembers against what the repository contains. The fix is manual: grok memory clear, then rerun it. If you take one operational habit from this article, take that pairing: git reset and grok memory clear go together, or your agent will keep acting on ghosts.

Worth naming the root cause, because it is bigger than Grok. Engine memory records the code as discussed, not the code as committed. That gap between conversation-truth and repo-truth exists in every tool-native memory we have tested; what we hit is a category of failure, bigger than any one bug report.

Run A session 3: old and new conventions coexisting in the native inbox

Run 2: the native layer plus repo files

Run B started differently from the first turn. Before writing any code, the agent read AGENTS.md and restated the protocol: read Active at session start, load Standby only when relevant, never rewrite Dormant. Then, unprompted, it anchored the commit convention into AGENTS.md itself, adding a ## Commits section. The rule moved from conversation into a file git can diff.

Session two made the same transliteration call, but the decision landed in MEMORY.md instead of a private store: one pointer line in Standby, one dated entry in Dormant carrying the full why (stripping the accent in "Café au lait" would yield caf-au-lait, missing the e entirely). A teammate cloning the repo gets the decision and its reasoning in the same checkout.

Session three is where the governance hooks earned their keep. Before touching code, the agent recited the rule that a canonical-check change requires a dated Dormant entry and a Verify-section update in the same commit. It then committed the backlog from sessions one and two first (462905d), and landed the convention change, the AGENTS.md Verify edit, the Dormant line, and the code together in one commit (b067b67). It also declared its own native-memory inbox out of bounds for this project: the repo files are the source of truth. When pytest turned out to be missing from PATH, it prefixed PATH for that one invocation and stated explicitly that it had changed nothing outside the project. Contrast Run A's silent pip install --user.

Run B session 3: the governance hook firing, backlog committed first, then one atomic convention commit

Session four ran the closeout protocol end to end: triage the memory tiers, drift check against git ("tree clean at b067b67"), confirm nothing new was worth recording, note process housekeeping. 38 seconds, and the day ended with a clean tree and every decision inside a commit.

Session five onboarded in about four minutes: read the Active tier, picked up every standing convention and prior decision, and implemented --max-len exactly the way the Dormant entry prescribed, backing up to the previous separator instead of cutting mid-token, 17 tests passing under pytest -q. And here is the honest part: it did not commit. Session three committed because a hook fired; session five had no hook for "finish by committing," so four files sat in the working tree. Discipline follows hook coverage, not good intentions. That is a design gap in our system, written down here because you will hit it too.

Run B session 5: cross-session handoff picking up conventions and the Dormant decision record

Run B session 2: MEMORY.md layering, a Standby pointer plus a dated Dormant entry with the full why

What each layer covers

Dimension Native ~/.grok memory With our three-layer files
What gets recorded, what gets missed Conventions, decisions, environment facts, all captured automatically (11 observations, 3 topic files). Commit discipline: missing. Four sessions of work never reached git. Rules anchored in AGENTS.md, decisions layered as Standby pointers plus dated Dormant whys. Also missed the day-two commit: hooks, not defaults.
Who wins a conflict Last instruction wins, no arbitration; old and new notes coexist in the same inbox, recall picks opaquely Explicit chain: AGENTS.md over MEMORY.md over engine notes; the native inbox was ruled out of bounds
Who catches a wrong memory Nobody. The phantom-memory incident needed a human to notice and run grok memory clear Drift checks compare memory against git at wrap-up and handoff; mismatches surface on the spot
Second person or agent onboarding Notes live in one private directory on one machine; same-machine recall works, anything else starts from zero Memory ships in the repo; a fresh session onboarded in about four minutes, decisions and whys included
Auditability Memory is markdown plus SQLite outside git; the log stopped at the first commit while work piled up Memory changes ride in the same commits as the code (b067b67); 100% traceable by closeout
Does memory survive switching tools No. The store is Grok Build's private format and path Yes. AGENTS.md is a convention most coding agents read; plain text in git, engine optional
Privacy and residency Local files, but invisible to the team and off by default; the docs' path is wrong Lives in the repo, travels with remotes; transparent, but a human must gate what goes in
Setup cost Effectively zero: two lines in a config file Low: drop in two files from the template repo and adjust one section

Bottom line: native memory is a solo superpower. File-based memory is team infrastructure. They stack.

Where each layer belongs

None of this makes Grok's native memory a bad feature. It is a good personal layer: zero configuration, automatic capture after every turn, recall that genuinely worked across sessions on day two. If you work solo on one repo with one tool, it covers most of what you need, and the official memory policy is notably transparent about what it leaves out: task state, tentative conclusions, secrets, and anything your docs already cover.

Our favorite outcome from here would be official support for the pairing. If Grok Build's memory ever learns to prefer repo-resident AGENTS.md and MEMORY.md files, and reconcile its own notes against them, the two layers stop running in parallel and become one system: native capture feeding governed, repo-side memory out of the box. We would adopt that on day one, and we suspect every team running more than one agent would too.

The file layer answers different questions, the ones that show up the moment a second person joins: can a teammate see what the agent decided, and why? Can you audit it in git? Does it survive switching engines? For us this is the third coding agent the same two files have run on. The engine changed; the memory didn't.

The two compose. Grok's notes keep doing their per-machine capture underneath, while the repo files govern what is allowed to become true for the team. And on residency: this is the same tool that faced scrutiny over repository uploads in July, shortly before open-sourcing the CLI and promising zero data retention. Where memory lives and who can read it is a question Grok Build's own history puts on the table. Notes in your repo, under your git host, your access controls, is one clean answer.

The honest limits

The 100% figure is taken at closeout, and day two exposes the gap: the handoff session left four files uncommitted because no hook told it to finish with a commit. Commit discipline in our system follows hook coverage, and coverage is incomplete. We are fixing that hook next, and the fix will be in the repo's changelog when it lands.

Sample size is one project, one machine, five tasks, one afternoon. It's a small sample, and we're upfront about that rather than overgeneralizing. Grok Build's memory is also three weeks old; the conflict-coexistence and phantom-memory behaviors may well improve, and the native store format is theirs to change.

Try it

If you want to reproduce this: enable memory with [memory] and enabled = true in ~/.grok/config.toml, expect the store at ~/.grok/memory-v2/, and pair every git reset with grok memory clear. Then drop the two files from wallaby-agent-rules into a copy of your project and run your own five tasks. The whole system is plain markdown: an AGENTS.md entry file, a three-tier MEMORY.md, and the write discipline that goes with them.

This series has the longer version of each piece: the three-file structure, the audit layer, the probation mode for letting rules earn their way in, and the Hermes memory audit, the previous engine in this experiment. Grok Build is the third. The engines change. The memory doesn't.

Your data, your business

Three commitments, verbatim from our privacy policy: No content logs. No training on your data. No usage reports built from your traffic. Usage lines record token counts, costs, and timing — never prompts, never completions.