What has to live in code
The gating page argued that code should own the threshold. This page goes one level stronger: some checks should not be asked of the model at all.
The strongest boundary statement in public
The hardest line comes from an agent-loop implementation:
A tool call with a high risk score forces human authorisation, and no probability can override it.
“No probability can override it” deserves to be pulled out. It does not say “confidence must be very high” — it says this rule is not in probability space. The distinction matters:
- “Auto-approve at ≥ 0.95” is a threshold rule — in principle some high enough probability would pass it.
- “High risk requires human approval” is a structural rule — true regardless of what the model says.
Those are not equally reliable. Your system should be explicit about which checks are which, and there should be as many structural rules as possible.
Run deterministic rules first, give the model only the grey area
Another phrasing of the same pattern:
Run deterministic rules first: whatever can be decided is blocked or allowed outright, and the model is only an optional backend. Routine commands are decided locally and never sent to the model at all.
Running rules first has an underrated benefit: it shrinks the attack surface. A judge that depends purely on the model behaves unpredictably under prompt injection; a judge that runs rules first confines injection to the small slice the rules cannot decide.
One published injection test reports 300 injection attempts and states plainly what the gate caught and what slipped through. That is far more useful than a bare claim of “injection protection”.
Secrets and sensitive data: block locally first
One secret detector lays the ordering out clearly:
| Case | Handling |
|---|---|
| Known secret formats | Blocked locally, never sent to the model |
| Unknown high-entropy strings | Masked first, then the masked text goes to the model |
| Model says ≥ 0.80 | Block |
| Model says ≥ 0.30 | Ask a human |
| Model unavailable | Ask a human anyway |
Two decisions to remember:
- Masking comes first. You cannot send something to an external service in order to find out whether it is a secret. The check itself leaks.
- A dead model still asks a human, rather than allowing. That is fail-closed in concrete form — service unavailability is not a reason to quietly lower the security bar.
Measured: 6/6 secrets blocked, 0/6 benign inputs falsely blocked.
Budget caps, signed receipts, audit
Some checks have nothing to do with what the model says, but have to live in the same system:
- One runtime gives dangerous-command blocking and budget caps to deterministic code, and records every verdict in an Ed25519-signed receipt chain.
- Another confines the risk score to a code-controlled 0–100 scale and signs every verdict with ES256 (540 calls, 99.76% and 0 false positives at a threshold of 65–75).
The point of signing is not to protect against the model — it is so that afterwards you can tell what happened. When an automated decision causes an incident, “what the model said and what the code did about it” has to be independently verifiable rather than a matter of trusting the logs.
Filter injection in two places
One guard layer’s design deserves its own mention because it names two easily missed spots:
Filter once before a tool call runs, and again before a tool result is read by the agent.
The second is the one people forget. Filter only the input and an attack can hide in what the tool returns — a fetched page, a file’s contents — which enters the context having bypassed the input-side check entirely.
A related rule from a retrieval-augmented implementation: never output a number that does not appear in the cited sources.
Sandboxing and permissions
Besides filtering content, the other layer is limiting blast radius: protected paths always go through human confirmation; agent execution runs in isolated workspaces (one implementation uses Docker); sub-agent permissions are separated from the parent’s.
A reference table
| Check | Which kind | Which way it falls on error |
|---|---|---|
| Known secret formats | Deterministic rule | Block |
| High-risk-operation authorisation | Structural rule | Deny (a human must approve) |
| Auto-approving read-only commands | Threshold rule | Fall back to the prompt |
| Injection inside tool output | Deterministic filter + model | Block |
| “Is the agent done?” | Efficiency gate | Allow (fail-open) |
| Format and style checks | Efficiency gate | Allow |
The middle column is the point: the closer to the top, the less the model should be involved.
In one sentence
A check that can be overridden by a probability eventually will be. Decide first which checks are not allowed to be overridden; only then tune thresholds.