Summary
As agents take on work that no person observes, the value of a task moves from its output to the account of how it was produced. Teams already have a well-studied mechanism for this, the debrief, and early work on language-model agents points the same way. We propose a definition of done for agentic work, argue that unexplained agent work should be treated as unfinished, and describe how closed tasks feed back into memory.
Work nobody watched
Teams remember collectively by knowing who knows what, which Wegner called transactive memory [1]. If a colleague paid an invoice, you can ask them whether they checked the purchase order. An agent breaks that chain. Unless it writes down what it did and why, nobody knows, including the next agent.
An old practice with a new reader
Closing notes are not new. The U.S. Army made the after-action review a routine part of training decades ago: what was supposed to happen, what actually happened, why, and what to do differently. Hospitals hold morbidity and mortality conferences to review what went wrong in patient care. Software teams write blameless postmortems after outages, a practice Google described in detail for its site reliability engineers [5].
Each of these was designed for human readers who were in the room or nearby. What changes with agents is the reader. The next reader of a closing note may be another agent, minutes later, deciding how to handle a similar task, and it will take the note literally.
The evidence for writing it down
Debriefs work. A meta-analysis by Tannenbaum and Cerasoli found that team and individual debriefs improve performance by roughly 25% on average over comparable groups without them [2]. Agents show a similar effect: in Reflexion, language-model agents that wrote short verbal reflections after each attempt, and kept them in memory, did better on later attempts [3].
Writing down what's next matters too. Masicampo and Baumeister found that making a specific plan for an unfinished goal reduced its intrusion on unrelated tasks [4]. A follow-up captured as a task frees attention; a follow-up buried in a paragraph doesn't.
Our definition of done
A task is done only when all four are true:
- 01The outcome is reached, or the task is explicitly dropped. Cancelled counts as done, with one line on why.
- 02An outcome note is posted as the closing comment, in four parts: what happened, decisions and why, links, follow-ups.
- 03Follow-ups exist as new tasks, not as sentences in the note.
- 04Files are filed: saved to their long-term home, not just linked from a closed task.
The fixed four-part structure is deliberate. Free-form summaries vary in length and focus; a uniform note can be skimmed in seconds and indexed reliably by the memory it eventually feeds.
Who writes the note
For work an agent did, the agent writes the note as its last step, with links to its own logs and files. For work a person did, the agent drafts the note from the activity it can see, such as emails sent, files changed and messages exchanged, and the person confirms or corrects it. Confirming a draft takes seconds. Writing from scratch is what makes people skip it.
The contrarian part
If a task can't be redone from its closing note, it isn't done.
Consider an agent that pays an invoice correctly but leaves no note. The outcome is right, but you can't tell whether it checked the amount against the order, whether the vendor's bank details changed, or whether the same invoice was paid last week. You can't audit it, repeat it or learn from it. For agentic work, a correct result with no account is a liability. We treat it as a failed task.
What a bad note looks like
Most bad notes fail in one of three recognizable ways:
- The empty note: "Done." It records that something happened and nothing about what.
- The diary: a long narrative of every step, with the one decision that mattered buried in the middle.
- The unsupported claim: "Verified the vendor details." Verified against what? Without a link to the record checked, the sentence can't be trusted, and a model can write it whether or not the check happened.
The four-part format is designed against all three. It forces a decision section, keeps the narrative short, and makes links a required part rather than a courtesy.
Scaling the note to the task
Not every task deserves the same note. We use three sizes:
- Trivial tasks, like reordering supplies, get one line that names the outcome and links the receipt.
- Routine tasks get all four parts, each a sentence or two.
- Consequential tasks, anything involving money, commitments made on someone's behalf, or decisions that are hard to reverse, get all four parts with evidence links for every claim, and a person reviews the note before the task closes.
The size is decided by rules, not by the agent's sense of importance, for the same reasons we give in our paper on auditable routing.
Then it goes back through the loop
A closed task re-enters the same loop as any new email. Most are simply cleared. Some notes are promoted into long-term memory. Some reveal new work. Every processed item is marked as handled, so a closed task can't recreate itself.
Notes are how the next agent learns
The Reflexion results suggest something beyond record-keeping. When a similar task arrives, say, the next invoice from the same vendor, the agent retrieves the last outcome note before acting: what was checked, what went wrong, what the person decided. Each note makes the next run of the same kind of task a little better, without retraining anything.
Where we might be wrong
- Overhead. Most tasks are trivial, and a four-part note for "reorder printer paper" is noise. We scale the note to the task: a trivial task gets one line.
- A wrong note is worse than no note. Language models can write a confident summary of work that didn't happen. Notes must link to evidence, such as logs, files and transaction records, and a note that can't be checked against them is flagged.
- Goodhart's law. Once notes are required, they risk becoming boilerplate that satisfies the checklist without saying anything. We test notes by use, not by presence.
Finally, notes concentrate sensitive information. A good note about a payment contains the amount, the vendor and the account checks performed. Notes inherit the access rules of the task they close, and they are filed only into stores the customer controls.
To avoid that trap, we don't score notes on length or completeness of headings. We score them on use: whether anyone, person or agent, retrieved the note later, and whether the redo test passed when it mattered.
What we're measuring
- Redo test: given only the closing note and access to the same tools, can a second agent reproduce the outcome? This is our working definition of a complete note.
- Note reuse: how often notes are retrieved later by people or agents.
- Evidence coverage: the share of claims in a note backed by a linked log, file or record.
References
- [1]Wegner, D. M. (1987). Transactive memory: A contemporary analysis of the group mind. In Theories of Group Behavior. Springer.
- [2]Tannenbaum, S. I., & Cerasoli, C. P. (2013). Do team and individual debriefs enhance performance? A meta-analysis. Human Factors, 55(1). Link ↗
- [3]Shinn, N., et al. (2023). Reflexion: Language agents with verbal reinforcement learning. NeurIPS 2023. Link ↗
- [4]Masicampo, E. J., & Baumeister, R. F. (2011). Consider it done! Plan making can eliminate the cognitive effects of unfulfilled goals. Journal of Personality and Social Psychology, 101(4).
- [5]Beyer, B., Jones, C., Petoff, J., & Murphy, N. R. (Eds.). (2016). Site Reliability Engineering: How Google Runs Production Systems. O'Reilly. Chapter 15: Postmortem culture. Link ↗