Confidence gating and escalation to a human
The previous page argued that code should own the threshold. This page is about what the threshold divides — the three bands on either side of it.
Three bands, not two
Real projects rarely make it “automatic or not”. Three bands is the common shape:
| Band | Condition | Action |
|---|---|---|
| Automatic | Confidence above the upper threshold and extra checks pass | Just do it |
| Escalate | Falls in the middle | Hand to a human |
| Drop | Confidence below the lower threshold | Do nothing, and do not spend a human’s attention either |
The third band is easy to overlook and is the cheapest of the three. “Not looked at” and “a person glances at it” differ by a human.
One triage implementation is typical: page only when P(SEV1) + P(SEV2) ≥ 0.80; drop
below 0.20 only when an actionability check also disagrees; send the middle band to
a human.
The middle band is where the actual problems live
Once the three bands exist, every hard question moves into the middle one:
It is human capacity, not spare capacity. The wider the band, the more escalations, the more a person handles. A 9% escalation rate sounds low — until the base rate is thousands a day.
It needs a fail-safe, or it gets silently stuck. Once you hand a decision to a human, nothing happens until a human looks. On a dashboard that looks exactly like “everything is fine”. The triage tool’s answer is blunt: no acknowledgement within 15 minutes, page anyway. Better to wake someone than let a P1 sit in a queue.
The escalation condition itself can fail. An autonomy gate inside an agent
continues only when a Choice clears 0.6 and an autonomy-safety Noul clears 0.5,
returning the turn to the human otherwise. Both numbers must pass — because “what should
I do” and “is this safe” are separate questions, and one confidence cannot honestly
answer both.
Setting the threshold: measure first
Picking 0.8 out of the air is the most common approach and the one most likely to become noise three months later. The reproducible approach is to measure, on your own labelled data, the threshold required to reach your target accuracy.
A public example: with 0.7 as the low-confidence gate, a run over n=130 caught 3 of 3 misclassifications at a 9% escalation rate. The sample is small. Its value is not the 0.7 — it is that this is how such a number should be reported: sample size, escalation rate, and what was caught.
Latency is the underrated risk
The triage tool’s measured numbers deserve their own table:
| Metric | Value |
|---|---|
| p50 | 418 ms |
| p95 | 1477 ms |
| Cost | about $0.04 per thousand |
| Slowest call | 151 ms short of its own 2-second timeout |
p95 is 3.5× p50. That means the tail matters more than the average — because timeout policy fires in the tail, and when it fires you are on the fallback branch, not the branch you designed. The slowest call came within 151 ms of the timeout, which is to say the system’s determinism depended on a coin that happened not to land badly.
So gating systems usually fail not by “judging wrong” but by timing out and taking the fallback. Test the fallback independently of the main path.
Two fallback directions, both correct
When something errors, public implementations split into two camps, and both are right:
- fail-closed: a permission judge denies on timeout or error. A safety gate has to work this way — a fail-open safety gate is a silent hole.
- fail-open: a Claude Code Stop hook allows on any error. An efficiency gate has to work this way — a fail-closed efficiency gate deadlocks the whole flow.
The criterion is not “which is safer” but which cost is larger: a false block or a missed block. And it has to be decided when you write the code, because what it determines is the exception path, not the main one.
In one sentence
In a three-band design the work is in the middle band: its width is a cost, its fail-safe is a reliability property. Design both explicitly rather than defaulting.