Watching an agent with a decision model
Coding agents are where these models get used most densely, and the uses cluster in three places. All three share one structure: take the judgement out of the generative model, give it to something cheaper and faster, and let code act on a typed answer.
1. Before a tool call: allow or not
The most direct spot. A Claude Code PreToolUse hook does this:
Ask a
Noul: is this shell command strictly read-only? Auto-approve at 0.95, otherwise fall back to the normal permission prompt — and never deny.
Three design decisions are worth naming:
- 0.95 is high. The bar for auto-approving sits well above the bar for “this looks risky”. A false allow costs a machine; a false block costs one extra confirmation. Asymmetric costs belong in asymmetric thresholds.
- It never denies. It only ever adds an approval; everything else returns to the original system. A gate that only adds capability is far easier to test — its worst case is having been no help.
- Dangerous commands never reach the model. A local hard-no list and an injection filter catch them first. This matters more than the threshold — see What has to live in code.
Measured result: 0 of 8 state-changing commands were auto-approved.
2. When the agent wants to stop: is it done?
Another frequent spot. A Stop hook asks the decision model whether the agent is finishing too early, by putting its plain-language completion rules to it.
The more restrained version works like this:
Spend one four-question call only when files changed and no passing check has run since. Fail open on any error.
“Only ask when it is worth asking” is the crux for this class — these hooks run every turn, and calling unconditionally multiplies cost by the turn count. The condition above is “unverified changes”, which is exactly the moment an agent is most likely to declare victory without having checked.
A related hook keeps an agent from finishing early by judging plain-language completion rules rather than trusting the agent’s own claim.
3. When context fills up: what is still needed?
The third spot is context compaction, where approaches diverge the most.
Traditional compaction summarises: hand a stretch of conversation to a model and get prose back. The original is gone, and summarisation is irreversible.
The decision-model approach does not rewrite — it only scores:
| Approach | What is preserved |
|---|---|
| Score each tool call and result for “still needed” | Lines judged to be kept stay verbatim |
Move low scorers into a store, leave an expand() pointer in place |
Nothing is deleted, only folded |
| An append-only frozen prefix | The prompt cache is never broken |
| Trim long Bash output before the model sees it | Terminal noise never enters the window |
The shared sentence: the decision model’s job here is a binary keep/drop judgement, not a rewrite. That is precisely what it is good at and what generative models are bad at — and “fold instead of delete” turns a wrong judgement from “information lost forever” into “one extra expand”.
The counterpoint: safety gates and efficiency gates fail in opposite directions
The most instructive contrast in this part of the ecosystem: given the same question “what happens on error”, projects chose opposite answers, and both are right.
- A permission judge: deny on error or timeout.
- A Stop hook: allow on any error.
The criterion is not “which is safer” but whether this gate guards risk or throughput. Gates that guard risk (permissions, secrets, dangerous commands) should rather block a good thing than let a bad one through; gates that guard throughput (completion checks, format checks) should rather let something through than deadlock the flow.
In one sentence
Agents are the most natural home for these models because they need a great many cheap judgements whose results are consumed by code immediately. But at every one of those points, decide first: which way should this gate fall when it breaks?