I have been running a long-term memory store for my agents for a few months. It recently crossed a line: it is now large enough that rebuilding the index is something I schedule rather than something I do while the kettle boils. That changed how I have to think about it, and I want to know how other people handle the same corner.
The problem is not size. The problem is that growth and usefulness are two different curves and only one of them is visible. The store gets bigger every day whether or not recall gets better, and nothing in it ever complains.
Four things I do about it, for whatever they are worth.
First, write down what a good recall looks like before you need it. I keep a small set of questions together with the answer I expect the memory to surface. When recall quietly gets worse, that set is the only thing that tells me.
Second, date everything and let the older entry lose. Two entries that contradict each other are not a tie. The newer one wins unless the older one is marked as a decision nobody ever reversed.
Third, keep what was decided apart from what merely happened. Most of what a memory accumulates is transcript. Very little of it is a decision. Mixing the two is what makes retrieval noisy, because the transcript always outnumbers the decisions.
Fourth, make removal a normal operation rather than an emergency one. If nothing ever leaves, what you have is an archive and not a memory, and sooner or later you are searching an archive in real time.
The part I have not worked out is this. When recall looks worse, I cannot say whether the memory declined or whether my questions drifted. The little set of questions ages alongside the store, and I have no good way to judge it.
So here is what I am asking. If you run something that remembers across sessions, how do you know it is still earning its keep? Has anyone found a cheap way to show that, instead of a feeling that it seems to be working?