Docs

Confidence gating and escalation to a human

The previous page argued that code should own the threshold. This page is about what the threshold divides — the three bands on either side of it.

Three bands, not two

Real projects rarely make it “automatic or not”. Three bands is the common shape:

Band Condition Action
Automatic Confidence above the upper threshold and extra checks pass Just do it
Escalate Falls in the middle Hand to a human
Drop Confidence below the lower threshold Do nothing, and do not spend a human’s attention either

The third band is easy to overlook and is the cheapest of the three. “Not looked at” and “a person glances at it” differ by a human.

One triage implementation is typical: page only when P(SEV1) + P(SEV2) ≥ 0.80; drop below 0.20 only when an actionability check also disagrees; send the middle band to a human.

The middle band is where the actual problems live

Once the three bands exist, every hard question moves into the middle one:

It is human capacity, not spare capacity. The wider the band, the more escalations, the more a person handles. A 9% escalation rate sounds low — until the base rate is thousands a day.

It needs a fail-safe, or it gets silently stuck. Once you hand a decision to a human, nothing happens until a human looks. On a dashboard that looks exactly like “everything is fine”. The triage tool’s answer is blunt: no acknowledgement within 15 minutes, page anyway. Better to wake someone than let a P1 sit in a queue.

The escalation condition itself can fail. An autonomy gate inside an agent continues only when a Choice clears 0.6 and an autonomy-safety Noul clears 0.5, returning the turn to the human otherwise. Both numbers must pass — because “what should I do” and “is this safe” are separate questions, and one confidence cannot honestly answer both.

Setting the threshold: measure first

Picking 0.8 out of the air is the most common approach and the one most likely to become noise three months later. The reproducible approach is to measure, on your own labelled data, the threshold required to reach your target accuracy.

A public example: with 0.7 as the low-confidence gate, a run over n=130 caught 3 of 3 misclassifications at a 9% escalation rate. The sample is small. Its value is not the 0.7 — it is that this is how such a number should be reported: sample size, escalation rate, and what was caught.

Latency is the underrated risk

The triage tool’s measured numbers deserve their own table:

Metric Value
p50 418 ms
p95 1477 ms
Cost about $0.04 per thousand
Slowest call 151 ms short of its own 2-second timeout

p95 is 3.5× p50. That means the tail matters more than the average — because timeout policy fires in the tail, and when it fires you are on the fallback branch, not the branch you designed. The slowest call came within 151 ms of the timeout, which is to say the system’s determinism depended on a coin that happened not to land badly.

So gating systems usually fail not by “judging wrong” but by timing out and taking the fallback. Test the fallback independently of the main path.

Two fallback directions, both correct

When something errors, public implementations split into two camps, and both are right:

  • fail-closed: a permission judge denies on timeout or error. A safety gate has to work this way — a fail-open safety gate is a silent hole.
  • fail-open: a Claude Code Stop hook allows on any error. An efficiency gate has to work this way — a fail-closed efficiency gate deadlocks the whole flow.

The criterion is not “which is safer” but which cost is larger: a false block or a missed block. And it has to be decided when you write the code, because what it determines is the exception path, not the main one.

In one sentence

In a three-band design the work is in the middle band: its width is a cost, its fail-safe is a reliability property. Design both explicitly rather than defaulting.