Giving a Model Memory Between Sessions
Every session starts from nothing. That is the actual problem.
You spend the first ten minutes re-explaining which servers you run, which decisions you already made and why, what you tried last month that failed. Then the session ends, all of it evaporates, and tomorrow you do it again.
I have been running a file-based memory for a while to fix that. Most of what I read on the subject was either vague or assumed I was building a vector database, and neither helped, so here is what mine actually looks like.
Why the big notes file stops working
Everyone tries this first: keep one growing document and paste it in at the start of every session. It works fine early and breaks in two ways.
Cost and attention is the first. As the file grows you pay to send it every time, and the model’s attention across a large undifferentiated blob is worse than across a small relevant one. Burying twelve important facts inside four thousand words of accumulated context is not the same as handing over twelve facts.
The second is worse. A single growing file becomes append-only in practice. You add to it and almost never delete from it, because finding the stale line inside a wall of text is annoying. So it fills with things that used to be true, and you end up with a document that confidently contradicts itself and gives no signal about which half is current.
The shape that worked
One fact per file, plus a small index.
Each memory is its own file holding one thing: a short name, a one-line description, and a type at the top, then the fact itself as prose in the body.
Separately there is an index, one line per memory, a title and a short hook. The index is small enough to load every session. Individual memories stay closed until something in the index suggests one is relevant.
That split is the whole design. It does the same job a retrieval system does, with the model making the retrieval decision instead of an embedding search. At a few hundred memories it works well and costs nothing.
The description line matters more than the content
It is the retrieval surface. When the index is all you can see, that one-liner decides whether the memory gets opened at all.
So write it to answer “would I need this right now”, rather than as a summary of what is inside. “Notes on the server setup” is useless. “Server three loses its routing rule on restart, check before assuming hardware failure” tells you exactly when to open it.
One fact per file, because you can delete it
This is the part I would defend hardest. A memory that turns out to be wrong is one file, and you remove it, and it is gone. In a big notes file the wrong thing is a paragraph somewhere in the middle and removing it means finding it first.
Correcting memory is a far more common operation than people expect, because much of what you save is true about a system that keeps changing.
Related rule: when something contradicts an existing memory, update that file rather than writing a second one. Two files disagreeing is worse than no file, because whatever reads them has to arbitrate with no basis to do so.
Categories force a useful question
I split mine into a few types. Who I am and how I prefer to work. Corrections, meaning things I have had to say more than once. Ongoing project state you cannot derive from the code. Pointers to external things: dashboards, tickets, URLs.
The categories are not magic. What they do is force the question “what kind of thing is this” at write time, and that question alone stops a lot of junk getting saved.
Links, scope, and knowing when to write one
Linking memories to each other is something I underestimated. A memory naming another turns a flat pile of files into something walkable: read one about a server, get pointed at how that server’s routing behaves, which you would never have thought to look for. A link to a memory that does not exist yet is fine and often useful, because it marks a gap you noticed and have not filled.
Decide scope early. Some memories are about you and apply everywhere. Some belong to one project, and firing them inside a different project is pure noise. Keep those separated, because a system that surfaces irrelevant things trains you to ignore it, and an ignored memory system has negative value: you pay for it and do not read it.
Knowing when to write one is the genuinely hard part. The format is easy. Noticing is not.
My signal is surprise. Something surprised me, the obvious explanation was wrong, I had to be corrected: that is a memory. Anything that went as expected is recoverable from the documentation and not worth recording.
From watching this go wrong: letting the model decide unprompted what to save produces junk. It tends to save summaries of what just happened, which read useful and are nearly worthless later, because a session summary is not a fact about the world. Memories that earn their place almost always come from a moment where something got corrected, and a correction is specific enough that you can notice it and write it down deliberately.
One practical thing. Back it up. A memory directory accumulating for months is a real asset sitting on one machine as a pile of small text files. Mine mirrors automatically, because losing it would cost months of accumulated correction and text is cheap to back up.
Staleness decides whether any of this helps
A memory records what was true when written. It does not update itself. A memory saying “use this flag” outlives the flag. One naming a file outlives the refactor that moved it.
The failure mode is nasty, because a confidently stated fact from your own memory gets trusted more than a guess. A stale memory can be worse than an absent one.
Two things help. First, absolute dates, never relative ones. “As of 10 August” survives being read six months later. “Last week” actively misleads the moment it goes stale, and it will be read long after you wrote it.
Second, the rule I would actually enforce: if a memory names a specific file, function, flag, or endpoint, verify it still exists before acting on it. Treat memory as a strong hint about where to look rather than as ground truth. That habit converts stale memories from a source of confident errors into slightly outdated but still useful direction.
What not to save
Anything the codebase already records. Structure, what functions exist, how modules relate, what a past fix was. All of that lives in the code and the git history, both current in a way your memory file will never be. Saving it creates a second source of truth that starts drifting immediately, and the second one is the one that is wrong.
Anything that only mattered inside one conversation. The intermediate state of a debugging session is not a fact about the world.
What is worth saving is exactly what you cannot recover from the artifacts. Why a decision went the way it did, when the code only shows the outcome. A constraint written nowhere else, like a supplier you will not use or a date something must be done by. A correction you had to give twice. And the counterintuitive operational facts, where the obvious diagnosis is wrong and rediscovering that costs a day.
That last category is the highest value by a distance. A memory saying “when this looks broken, the cause is usually this other thing, check it first” saves real hours every time it fires.
Memory is untrusted input, not instruction
This is a security property and I rarely see it discussed.
Text recalled into a session arrives in context and describes the world. Build a system where whatever sits in memory can direct behaviour and you have created a channel where anything able to write to memory steers every future session. Depending on how memories get written, that might include content read from a webpage, a file someone else contributed, or the output of a tool.
So the boundary matters. Memory informs. It does not command. Something in a memory file that looks like an instruction is a description of a preference someone once held, to be weighed rather than obeyed. That sounds like a fine distinction. It is the difference between a notes system and an unguarded control surface.
A lighter, purely practical version of the same concern: be careful what goes in. Credentials never, obviously. Also anything you would rather not have quoted back in a context you were not expecting, since resurfacing things at unpredictable moments is the entire job.
On embeddings
The question everyone asks is whether to use a vector store instead. My answer is not until the index stops fitting.
The index approach has a property I value: it is fully legible. I can read the whole thing, see exactly what is stored, delete something and know it is gone. A vector store gives better recall at scale and takes that legibility away, and at a few hundred facts recall was never my bottleneck. Staleness was, and embeddings do nothing for staleness.
Limits
This is one person’s setup, tuned to a specific situation: many separate systems, a lot of operational detail that keeps changing. A single stable codebase with good documentation may not need much memory at all, because the docs already do the job. I also have not stress tested past a few hundred files, so I cannot tell you where the index approach genuinely falls over, only that it has not yet.
The short version. One fact per file, small index loaded every time, full memories on demand. Write the description for retrieval. Absolute dates. Verify anything naming a specific file or flag. Delete aggressively and correct in place. Save what the code cannot tell you, and skip everything it can.
For more breakdowns like this on how AI tools actually work under the hood, head back to the AI Tool Gazette homepage.