Summary
The prevailing bet in personal AI is total recall: capture every email, meeting and chat, embed it, and retrieve it on demand. We think this is backwards for most of what people receive. Work on human memory suggests forgetting is adaptive, studies of email suggest people rarely go back for what they file, and experiments with language models show that irrelevant material makes answers worse. We describe the routing loop we use, state our position, give the evidence and the strongest objections, and list what we are measuring to find out if we are wrong.
The total-recall bet
The idea is old. Microsoft Research's MyLifeBits project set out to store everything a person sees and hears [1]. Cheap storage and embeddings have revived it: if keeping everything costs almost nothing, why decide? Our answer is that storage was never the expensive part. Attention is, and so is the precision of whatever has to search the pile later, whether that is a person or a model.
The dream is older than computers. In 1945 Vannevar Bush imagined the memex, a desk that would store a person's books, records and correspondence so they could be consulted "with exceeding speed and flexibility" [13]. Sixty years later, Gordon Bell and Jim Gemmell described living it in their book Total Recall [14]. What both visions emphasized was capture. What they underplayed was the cost of finding the right thing again, and of being misled by the wrong one.
Three lines of evidence
First, human memory forgets on purpose. Anderson and Schooler showed that how available a memory is closely tracks how likely it is to be needed, given the statistics of the environment [2]. Forgetting is not a defect; it is the memory system betting on what will matter. Complementary learning systems theory adds that a fast, temporary store feeds a slow, durable one through selective consolidation [3]. Not everything is meant to make it across.
Second, people rarely refind, and filing doesn't help much. In a study of how people refind email, Whittaker and colleagues found that those who carefully filed messages into folders were no more successful at finding them later than those who relied on search, and spent extra time filing [4]. Bergman and colleagues complicate the picture: even with better search engines, people still prefer to navigate to things by location [5]. Both findings matter to us. The first says curation by hand is a poor use of time. The second says people want to know where things are.
Third, more context can make models worse. Shi and colleagues showed that adding irrelevant information to a problem sharply lowers the accuracy of large language models [6]. Liu and colleagues found that models use information at the start and end of a long context much better than information in the middle [7]. A memory stuffed with everything doesn't just make retrieval slower. It makes the answers built on it less reliable.
Attention is the scarce resource
Storage keeps getting cheaper. Reading doesn't. Every item kept in the place an agent searches by default competes for two limited things: the model's context window and the person's attention when they review the answer. Context windows have grown enormously, but the evidence above suggests that a model's ability to use what is in its window does not grow with it.
There is also a compounding effect. Whatever sits in curated memory is a candidate for every future question. One stale price list or one superseded decision doesn't cause one bad answer; it can quietly contaminate every answer that touches the topic for years. Curated memory behaves less like a filing cabinet and more like a shared water supply.
Two tiers, not one pile
Saving less does not mean deleting more. We separate two things that "memory" usually blurs:
- Raw archives: the email archive, the drive, the chat history. Cheap, untouched, searchable when you go looking. We leave these alone.
- Curated memory: the small set of notes, decisions and documents an agent retrieves from by default when it acts for you. This is where the bar is high.
The valuable step is promotion: deciding what, out of everything that arrives, earns a place in curated memory. That is where our loop spends its effort.
The loop
Every item from every inbox goes through the same questions, in this order. The structure borrows openly from David Allen's split between actionable items and reference [8]; what's new is that an agent asks the questions, on a schedule, by written rules.
- 01Is it actionable now or soon? It becomes a calendar event, a task for an agent or a task for a person.
- 02If not, is it valuable? If not, it's cleared from the inbox and left in its raw archive.
- 03If it's valuable, will it inevitably become actionable later? It goes on a someday list.
- 04If not, it's promoted into curated memory, in the store that fits it.
The default answer to question two is no. When in doubt, nothing is promoted.
What promotion looks like in practice
- A flight confirmation becomes a calendar event. The email itself is cleared and stays in the email archive.
- A newsletter is cleared. Nothing is promoted.
- A supplier's new price list is promoted into documents, and the old one is marked as replaced so it stops answering questions.
- A long chat with a contractor in which a decision was made is summarized. The summary is promoted; the transcript stays in the raw chat history.
The pattern is the same each time: promote the smallest thing that preserves the decision or the fact, and leave the bulk where it already lives.
Working memory and long-term memory
Human working memory is small and temporary [9], with a capacity of only a handful of items [10]. We mirror that on purpose: tasks, the calendar and the someday list are working memory, and they are meant to empty out. MemGPT draws a similar line for language-model agents, between a small main context and a large archival store [11]. Finished tasks are where consolidation happens: like the periodic reflections in Park and colleagues' generative agents [12], a closed task is summarized, and only the summary is considered for long-term memory.
A memory system should be judged by what it lets you forget, not by what it keeps.
What this means for product design
Saving less is a default, not a ban. Three design choices follow from it.
- 01Promotion is visible. Once a week, the person sees what was promoted into curated memory and can demote anything with one click. Bergman's finding that people want to know where things are applies to memory as much as to files.
- 02Demotion is cheap and common. Facts go stale: prices change, decisions are reversed, people move on. Curated memory needs an expiry habit, not just an intake habit.
- 03Agents cite where an answer came from. When an answer draws on curated memory, it links to the note or document it used, so a stale source is caught the first time it misleads.
None of this requires a smarter model. It requires deciding, in writing, that a smaller memory you trust beats a larger one you have to double-check.
Forgetting, written down
If forgetting is a feature, it needs rules like any other feature. Ours are simple. A document that is replaced, like a price list or a policy, retires its predecessor the moment the new one is promoted. A decision that is reversed keeps its history but stops answering questions about the present. Notes about people and vendors are reviewed once a year and demoted if nothing has touched them. And anything promoted but never retrieved within a year is flagged for review rather than kept by default.
These rules are deliberately boring. The point is not to be clever about what to forget; it is to make forgetting routine, so curated memory stays small enough that a person could, in principle, read all of it.
Where we might be wrong
- Value is often obvious only in hindsight: the receipt you need for a warranty claim two years later. Our answer is the raw archive, which still has it. But if refinding from raw archives turns out to be routinely painful, the case for aggressive curation weakens.
- Models are getting better at long contexts. If tolerance for irrelevant material improves enough, the retrieval argument for saving less loses force, and only the attention argument remains.
- People like knowing that everything is kept. Bergman's results suggest the feeling of control matters, not only retrieval speed. A system that saves less has to show its work to earn the same trust.
One more risk deserves its own line: the promotion gate is itself a model judgment, and it can be wrong in systematic ways, for example by favoring whatever is recent or emotionally loud. That is why the gate is governed by written rules and logged decisions, as described in our paper on auditable routing.
What we're measuring
We have not yet run these measurements at scale; this is the plan we will report against.
- Refind rate: the share of promoted items retrieved at least once within 90 days. If it is high, our bar is right or too high. If it is low, we're still saving too much.
- Precision as memory grows: answer quality on a fixed set of questions as curated memory goes from hundreds to tens of thousands of items.
- Regret rate: how often someone needs something that wasn't promoted, and whether search of the raw archive recovered it.
References
- [1]Gemmell, J., Bell, G., & Lueder, R. (2006). MyLifeBits: A personal database for everything. Communications of the ACM, 49(1).
- [2]Anderson, J. R., & Schooler, L. J. (1991). Reflections of the environment in memory. Psychological Science, 2(6).
- [3]McClelland, J. L., McNaughton, B. L., & O'Reilly, R. C. (1995). Why there are complementary learning systems in the hippocampus and neocortex. Psychological Review, 102(3).
- [4]Whittaker, S., Matthews, T., Cerruti, J., Badenes, H., & Tang, J. (2011). Am I wasting my time organizing email? A study of email refinding. CHI 2011.
- [5]Bergman, O., Beyth-Marom, R., Nachmias, R., Gradovitch, N., & Whittaker, S. (2008). Improved search engines and navigation preference in personal information management. ACM Transactions on Information Systems, 26(4).
- [6]Shi, F., et al. (2023). Large language models can be easily distracted by irrelevant context. ICML 2023. Link ↗
- [7]Liu, N. F., et al. (2024). Lost in the middle: How language models use long contexts. Transactions of the ACL. Link ↗
- [8]Allen, D. (2001). Getting Things Done: The Art of Stress-Free Productivity. Viking.
- [9]Baddeley, A. D., & Hitch, G. (1974). Working memory. Psychology of Learning and Motivation, 8.
- [10]Cowan, N. (2001). The magical number 4 in short-term memory. Behavioral and Brain Sciences, 24(1).
- [11]Packer, C., et al. (2023). MemGPT: Towards LLMs as operating systems. Link ↗
- [12]Park, J. S., et al. (2023). Generative agents: Interactive simulacra of human behavior. UIST 2023. Link ↗
- [13]Bush, V. (1945). As we may think. The Atlantic Monthly.
- [14]Bell, G., & Gemmell, J. (2009). Total Recall: How the E-Memory Revolution Will Change Everything. Dutton.