Summary
An agent's workflow can be fixed while every decision inside it stays probabilistic. The usual safeguard is to ask the model how confident it is and escalate to a person when confidence is low. We argue that self-reported confidence is the wrong signal, and that human approval degrades quietly when it's asked for too often. We propose four practices (rules first, disagreement-based escalation, an escalation budget and ordered questions) and one metric, the judgment surface, which should shrink over time.
Fixed structure, probabilistic nodes
A routing tree looks deterministic: is it actionable, is it a calendar event, can an agent do it? The shape never changes. But each question is answered by a model, and the same item can land in different places on different days. Lowering sampling temperature doesn't fully solve this; even at temperature zero, inference can vary between runs because of how requests are batched on the server [1].
Determinism is a spectrum
It helps to stop thinking of an agent as either deterministic or not. Each decision can sit at one of four levels, from most to least predictable:
- 01Code: a fixed function with no model involved, such as checking whether an email has a calendar attachment.
- 02Written rules: plain-language conditions a person approved, applied by code or by a model that is only asked whether the condition holds.
- 03Constrained judgment: a model choosing from a short, fixed list of destinations, with structured output it can't step outside of.
- 04Open judgment: a model writing a free-form answer or plan.
Good agent design pushes each decision as far up that list as it can go. Routing should never be open judgment: the model picks one of a known set of destinations, and anything else is treated as a failure, not a creative answer.
Self-reported confidence is a weak signal
The obvious fix is to ask the model whether it's sure. The evidence is discouraging. Xiong and colleagues found that confidence stated in words by language models is systematically overconfident [2]. Kadavath and colleagues showed that models can estimate the probability their own answers are true reasonably well in some settings [3], but the GPT-4 technical report found that calibration got worse after post-training, the step that makes models helpful in chat [4]. The models people actually use are the post-trained ones.
We use disagreement instead. Self-consistency, sampling several answers and taking the majority, improves reasoning accuracy [5]; we use the same idea for safety. Each item not covered by a rule is routed several times. When the routes agree, the item proceeds. When they disagree, it goes to a person. This costs more compute, but only on the minority of items that rules don't already handle.
The extra cost is real but bounded. Repeated routing applies only to items no rule covers, and as rules accumulate, that share falls. The compute spent on disagreement is, in effect, the price of finding out where the rules are missing.
Humans in the loop fail quietly
"A person approves anything important" sounds like the end of the safety story. It isn't. Decades of research on automation bias and complacency show that when automation is usually right, people stop checking it [6]. An agent that escalates constantly is training its supervisor to approve without reading.
So we set an escalation budget: a cap on how many items per day reach a person. If the budget is exceeded, that's treated as a defect in the rules, not a reason to ask the person to try harder.
A worked example
A client writes: "Can we move Thursday's install to next week?" A naive rule sees a date and sends it to the calendar. But nothing should go on the calendar yet; someone has to check the crew's availability and reply. Routed three times, the model sends it twice to a person's task list and once to the calendar. The disagreement escalates it.
The person handles it and the log records why. After the same pattern appears a few dozen times with the same resolution, it becomes a written rule: reschedule requests become a task for a person plus a tentative calendar hold. The next hundred such emails never reach the model at all.
Setting an escalation budget
The budget should start from what one person can review attentively in a day, not from how many items the agent produces. For a small business owner that may be ten or twenty items, read properly, rather than a hundred skimmed. When the queue overflows, there are only three honest responses: write more rules, narrow what the agent is responsible for, or add reviewers. Lowering the bar for what counts as "reviewed" is not one of them.
The budget also makes the agent's quality visible. A falling escalation count with a steady override rate means the rules are working. A falling count with a rising override rate means problems are slipping through unreviewed.
The order of questions is policy
While designing our own loop, we first asked "Can the AI agent do it?" before "Is it a calendar event?" Since an agent can always add a calendar entry, every calendar item was sent off as a generic agent task and never reached the calendar. Nothing was wrong with any single decision; the order was wrong. We now put questions with clear, cheap answers first, so each question narrows what the next one has to decide.
The judgment surface
We define the judgment surface as the share of decisions made by model judgment rather than by written rules. Every decision is logged with its inputs and outcome. When the logs show the model making the same call the same way, again and again, a person turns that call into a rule. The judgment surface shrinks, and what's left for the model are the hard cases, which are also the ones a person should see.
Measure an agent by how small its judgment surface gets, not by how smart its judgments are.
This runs against the industry's instinct to measure agents by capability. For decisions about people's money, families and businesses, we think explainability is worth more than the last few points of accuracy.
Why not just fine-tune?
A natural alternative is to fine-tune a model on each customer's past decisions. We don't, for two reasons. It would mean training on private data, which we have promised not to do. And a fine-tuned model is a black box: it may make better calls, but it can't tell a person which rule it followed. A written rule can be read, questioned and deleted.
What a useful log entry contains
"Log every decision" is easy to say and easy to do badly. A log is only useful if a person can reconstruct a decision from it without re-running anything. Each routing entry records:
- The item, by a stable identifier, and where it came from.
- Whether a rule decided it, and if so which rule and which version of it.
- If the model decided: each sampled answer, whether they agreed, and the short reason given for each.
- The final destination, and whether a person approved, changed or reversed it, and how long they took.
- The model and prompt versions in use, so behavior can be compared across upgrades.
With entries like these, the questions that matter can be answered by querying the log: which rules fire most, where the model disagrees with itself, which decisions people reverse, and whether a model upgrade quietly changed behavior.
Where we might be wrong
The strongest objection is history. Rule-based spam filters lost to statistical ones [7], and Sutton's "bitter lesson" argues that general learning methods beat hand-written knowledge in the long run [8]. Rules are brittle, they pile up, and they fail on cases nobody imagined.
Our reply has two parts. First, our rules aren't invented up front; they're distilled from the model's own logged decisions, so learning still drives them, and a person signs off on each one. Second, the bitter lesson is about maximizing performance. Our goal is different: decisions that can be explained, reversed and audited. If a learned method can offer the same guarantees, we'll switch, and we'll say so.
What we're measuring
- Judgment surface over time, per Island: it should fall week over week.
- Override rate: how often people reverse a routing decision, split by rule and by model judgment.
- Rubber-stamp signals: approval time and approval rate. Approvals that take a second or two, with almost nothing ever rejected, suggest the person has stopped reading.
- Disagreement rate: how often repeated routes disagree, and how often the person sides with each.
References
- [1]He, H., & Thinking Machines Lab. (2025). Defeating nondeterminism in LLM inference. Thinking Machines Lab: Connectionism. Link ↗
- [2]Xiong, M., et al. (2024). Can LLMs express their uncertainty? An empirical evaluation of confidence elicitation in LLMs. ICLR 2024. Link ↗
- [3]Kadavath, S., et al. (2022). Language models (mostly) know what they know. Link ↗
- [4]OpenAI. (2023). GPT-4 technical report. Link ↗
- [5]Wang, X., et al. (2023). Self-consistency improves chain of thought reasoning in language models. ICLR 2023. Link ↗
- [6]Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3).
- [7]Graham, P. (2002). A plan for spam. Link ↗
- [8]Sutton, R. (2019). The bitter lesson. Link ↗