Agent memory as a file format
calpaterson.com
calpaterson.com
I find once there is one poisoned line of text it negatively affects everything else downstream. Instead, I use a temp/ folder with documents and use different files for different agents and models. Then I have to constantly prune and delete the files. Any information that can be extrapolated is just noise which negatively affects the agent. If you have a definition of a database structure and it has been implemented, that information should not be contained in any text document -- it is noise, will drift, and be impossible to debug why the agent keeps producing undesired behavior.
I have a ~/Projects folder. For example, I use Playwright with Chrome DevTools Protocol in order to do performance testing and leak detection. There is a script that handles this. My prompt is "Search ~/Projects for perf testing with CDP and Playwright and implement here". Point being, if I need anything I point to a resource or ask to search a resource and it will find it quick and, most importantly, tends to improve it every iteration.
If I was in an institution, I would have a repository and would rather just point the resource and say use that than have memory of it locally.
I turn it off for local agents because I bounce around a few and I really don't like the mostly implicit nature of it. I want to write my instructions in version control if I have anything to say consistently to an agent.
Just to start with: memoryfields are possible to use in a server/client system. That was a key aim and I do already use them over Amazon S3 (though not always).
I started, like you did, with a personal library of prompts. But the issue is that as your library of little pieces of prompts increases a) you get tired of constantly editing them yourself b) you have no easy way to export and share them with others c) it's frustrating that the agent doesn't "automatically" find your little bit of prompt on X even when clearly it is relevant - hence sem search.
I think a lot of people are still using the "personal library of bits of prompt" model. It is ok. But I wanted to propose an minimal, interchangeable standard for sharing them. So the idea of being an institution and having a shared memoryfield: that's something I want as well!
The spec, feedback greatly welcome:
https://github.com/calpaterson/memoryfield-spec/blob/main/SP...
I would like to hear more details of what it ended up with for a structure
Within projects, I make heavy use of path-scoped rules to intentionally bring context where it's needed, and also make heavy use of temp directories. LLMs are more than happy to produce ad-hoc memories/summaries/context docs that I can then point a session to, but I can be selective and intentional about it.
I like that memoryfield is portable, intentional and composable. I'm not convinced that sharing memoryfields between users will be practical, but I keep isolated virtual environments for absolutely everything. I like the idea of being able to intentionally bring collections of managed context around with me. There are other ways to do that, but will keep an eye on this.
You have to constantly tend the garden and weed these things out.
Whenever Claude makes some error ask it where and why? And then dig out the weed.
And of course if it’s in the context, probably time for a handover doc (which will need weeding) and started fresh.
There's a simple fix for this: do not let it write comments. Ever. This also has the nice property that there's much less AI slop to clean up afterwards.
My method is seven layers of files, administered differently: meta-knowledge, project seed, LLM wiki, code, tickets and todos, chat logs, artifactory. Ordered idea-to-reality. All git repos. That outgrows any context window pretty soon. My way out of that trap is to use links, both wiki links and git permalinks.
It keeps documentation from going out of date by embedding re-runnable verification checks directly inside markdown files (that the agents use, not humans). I use it with a handoff workflow to force Claude/GPT to re-verify facts before handing off documentation to the next session.
That being said, in coding over a longer time horizon, having the agent continually re-derive decisions/laws/facts/etc from your code is wasteful of tokens and time, and if your code doesn’t consistently apply them you can’t know the agent will make the correct choices.
You need memory of these important facts to avoid this expense or potential incorrectness. Memory does not itself scale though, without maintenance and pruning, and that has its own impacts on cost and correctness like the Chat memory.
“Damned if you do, damned if you don’t” at least until the agent can itself maintain its memory accurately - or some other non-human effort can achieve that.
e.g. docs/decisions/README.md (index with a blurb about each decision), docs/decisions/01-some-lesson.md (some architectural decision/pattern that you or the agents discovered).
ADR files have important sections like "rejected solutions" and "acceptable risks", and they're live files that can be refined and pivoted over time or retired to docs/decisions/archive/.
It's also nice to give each top-level bullet point some stable ID like "R1" for rejected solution #1, I3 for invariant #3. Agents use this stuff intelligently all the time like "This could be a time to reconsider D4/R2" = ADR #4, rejected solution #2.
The essential part being that your system ratchets into increasingly better decisions and invariants over time, and there's a place to actually put this stuff.
It's essential for automating high-quality software and something we couldn't be arsed to do much less update before AI.
Two agent.MD files that are very small. One on each project. One at parent project level.
Did 30M tokens through glm 5.3 flash today for 52c
Using pi and a few extensions my initial context is always 4k max
- memory systems are a specific type of knowledge base where you generate all the documents. You might as well generate them to be less than your embedding token limit to obviate the need for chunking.
- embedding models are getting better and are no longer just semantic averaging.
- small models are getting dirt cheap, making parallel reads cost manageable
What they describe is sort of the simplest architecture that takes advantage of these observations. I believe them when they say it works well.
I do suspect though that things like keyword lookup will completely fail if every memory is just a vector. Hence why something like Typesense hybrid search can still be useful.
I do think that having a set of token that are highly personalized to your project and to way you work is beneficial. I also think that the idea that this set of token will be constructed in the background without any work from the user is really appealing. So it's understandable that the 'memory' analogy became so popular.
But in my experience having a really good AGENTS.md file almost always produce better results than enabling memory.
Maybe we should start to think about how we 'train'/'onboard' agents into our projects, in a similar way that we do for new co-workers. Imagine if we could send the agent to our repo and ask it to learn our patterns and in the end we could quiz the agent to gauge how much it actually understood the project. Once he 'understands' the project we can start to use it to help with development.
In a very small scale (example, individual new features) I will sometimes ask the agent to explain me how things work (even though I already know how it works) so I can 'prime' the agent context with good data before starting any real work. But I'm not sure if this approach could be reliably scaled to work with any repo for any kind of work.
Works pretty well for "non-permanent" instructions (you don't want to put this info into your committed markdowns).
My main motivation was simply to save tokens, but it actually worked really well and improved speed as well.
One way I've toyed with a graph outline with it using "whitening" (https://arxiv.org/pdf/2104.01767v3) for embeddings rather than just text, so you add things like file path, nearest title/method/const etc. You have to have dummy text though because it fails with "null"; all embeddings need to carry some kind of text and of the same size.
so when the agent rembembers something, the memory would be an embedding that includes where they found the file, what method or const or whatever they're in, what the task they're working on is, etc. That all becomes a single embedding. You could imagine a metadata tag that also describes the tools they're using etc.
Map out a complete space of tags for whitening an embedding and there's surely a proper mix. Then when you're searching for things in the embedding, you also store some of the other metadata as plane strings & edges, which gets you some useful granularity.
Memory systems can also let AI load memories on demand so it's actually analogous to maintaining AGENTS.md plus twenty different "read this if you need to do X" markdown files.
I think the name "memory" makes sense given that the AI system is one singular system with central context (rather than a software org of multiple distinct humans with distinct memories), so the equivalent of institutional knowledge in the software org really is just akin to memory for the AI system.
But it's still text. You show people how the sausage is made and either they're confused or they're horrified.
The markdown explanation I was also unimpressed with, but RAG over hyperlinks is convincing to me.
---
title: Carbon Fibre Woks
created: '2026-03-01T09:00:00Z'
updated: '2026-08-22T14:30:00Z'
uuid: 6aa615f0-486f-48a7-a210-ba4f5ff18c8b
summary: Thermal properties of carbon fibre cookware
---
Carbon fibre woks conduct heat evenly, but...
So basically recfiles[1], and an entire suite of tools replaced by fopen and sed. Anyone who doesn't know of recfiles is doomed to reimplement it(poorly).This leads to, if useful at all, to this process of ad-hoc recall inference which has to happen in time. By itself this is not a problem.
The problem is that the model has to be constantly injected in context with the newest version of the "memory state" at each turn or relevant turn.
The newest coherent memory state is also a problem. More or less 50 years of not failures, but not success either. This may be even deeper problem than the externalization problem.
I think we are just in the very beginning and we are slapping database stuff to the transformer hoping it will work, but these deep neural-net architectures categorically show that they are not databases.
There will be a synergistic middle ground, but its shape is still not clear.
I wonder how much my system needs something like this. Between the invisible system memory of my random chats with Gippity, my Matt Pocock skills saving terminology and plans, and whatever else Cursor and Codex do, I don't think I feel a need for more agent memory. I do like how it's exposed and searchable, and not invisible. But I honestly just send my questions/tasks away to my magic agent and eventually it gets it right anyway; do I need more discrete memory my team has to maintain? (That's an earnest question, not disregard for this)
I.e. for the web we did that with link count etc.
We need some other mechanism for judging and ranking pieces of "memory" for "agents"
Seriously, isn’t this the core premise of RAG / modern embedding search systems?
That said, I think this article's basic idea of "text files + semantic search index" is a good way to implement memory because AI is already good at search, and semantic search is far more flexible than a knowledge graph and decreases the chance of the agent simply not being able to find a memory (or inserting duplicate memories, etc.) due to deciding to go into a slightly different branch than the one it needed to travel.
That said, I'd probably layer additional steps on top for further efficiency improvements. For example, automatic concise summaries of memory items, other supplemental indexing methods for better memory recall, etc. My ideal solution would be multi-step and therefore wouldn't solve the knowledge-graph's slowness, but it would solve the rigidness problem of a knowledge graph.
Also seems like a requirement for any sort of continual learning capabilities as well.
I am on linggen.dev , that is the one make daily agent work easier.
My main gripe so far is that I have to push the agent to maintain an organized graph. Once its there, it can be pretty useful to keep track of small notes which would otherwise be sprinkles all over my projects. The kg has some design limits that make poisoning hard to diagnose, so that's still on the list to fix.
And I agree: it's not a very special thing. That's why I propose: Markdown + a simple embedding.
That's agentic memory.
Now extend that to the rest of whatever system you’re working on. Do you want to manually have some preamble for every task you’re doing?
Even purely for convenience and dev experience it’s worth doing imo.
This buys you incremental writes, commit hash pinning and diffs for free. In addition to having a git archive outputting a zip export as well?
Main wrinkle is you would need to gitignore the sqlite database as it doesn't store very well in git (binary, changes lots per insert). But it's easy to regenerate anyway as a rebuild-able cache.
Git is supported in the spec, though the tool doesn't yet handle it (soon! git is nice as you get a log "for free" which helps agents understand more about the memories).
As for sqlite in git: yes you can gitignore it, you could use git-lfs, you could use an out of band database. It would be great if there were another format more amenable to git to store vectors in. I looked at csv closely, but I was worried about float formatting/representational bugs. There is a gap in the market for a format here.
Well regarding how CSV didn't quite work for you. Have you checked out recutils?
Its a GNU Project and at least does basic relational database operations on plain text files. You can easily install it in most linux package managers.
https://iwe.md/blog/your-agent-hates-walking-your-knowledge-...
- Store session turns in an sqlite-vec
- Provide the agent with an mcp to search the vec-db
- Let the agent write notes in md files along with an index / frontmatter
Along with the commit history, the vec-db gives the agent long-term memory. The notes allow the user to correct accumulation of false lessons.
Simpler but better.
Of course, ideally your data would be structured, but the agents will mostly be grepping anyway, and maybe look at sibbling files.
I just had a thought about similar thing - how to track human decisions on the codebase? Consider you are writing code together with AI, how you understand which code change happened because human asked for it
Should related new info update an existing page or crate a new one?
Incredible to watch things come full circle. Next thing you know, someone is going to figure out a binary encoding.
What people are exploring are other options as far as I can tell.
What you are offering is “don’t do that, this already works”. Which I guess is fine, but apparently not everyone is fully satisfied with the current generation of tooling.
Also fwiw grep is pretty poorly suited to semantic search and will only return the most basic of matches.
If you are really trying to build a useful memory search tool there are much better options than plain text search.
Of course a real search system with ranking and whatnot would be better, but you’re paying a different cost there.
I’m personally more interested in how the memory files get created and updated and generally managed, retrieving the correct ones doesn’t seem to be a big problem at the moment.
I primarily use gh copilot and they are by default hidden from the user. It’s also not great that they aren’t in the repo and every contributor has their own set of memories of different freshness, likely conflicting.
Both can be true: - It's useful to anthropomorphize agents when predicting behavior and - we have to use specific language to specify what we mean.
What does the author mean by "confuse the models" ? Are they talking about not picking right information? Picking the wrong information? Losing their previous context / task?
Part of setting up a proper eval is also deciding what we actually mean ourself. What are we actually optimizing for? It's not, e.g. % confusion, %rubbish, etc.
The article does point to it: retrieval latency, accuracy, etc.
> (optional) YAML frontmatter and
> (optional) SQLite vector index for semantic search
This is basically exactly what I use in a MCP service I built and it works pretty well. Can be enriched further if you use a storage system like S3 and take advantage of metadata.
"harness managed" memory is utter garbage, I am convinced, and I disable it immediately. The major problem being over time it degrades and sneaks in conflicting or outright false information. Then one day you'll swear it's drunk, and every time I got to this state and investigated, auto managed memory was always the problem.
> This is a common fear with memory systems but doesn't really apply to memoryfields. Irrelevant material is simply never surfaced by the semantic search.
This is so wrong. The Achilles' heel of this approach is the RAG. What makes it worse is having lots of memories that are outdated, wrong, hallucinated, or irrelevant.
Nothing beats curated data. Memory should be regularly reviewed, compacted, and cleaned up if it's no longer valid.