Docs

Cascades and parallel questions

The motivation behind both patterns is simple: a large model call is expensive, and most requests are easy. A decision call costs one to two orders of magnitude less than a generation, so “ask the cheap one first, escalate when unsure” is the obvious shape.

Cascades: only decide where you are sure

A head-to-head evaluation on medical hallucination detection gives a clean number:

Let the decision model settle the 37% of cases where it is at least 90% confident, and it preserves each large model’s own accuracy while removing 37% of the large-model calls.

This is worth quoting because it says what a cascade’s saving actually is: the share of traffic that falls inside your high-confidence region, not some fixed percentage. Tighter confidence means more saved — as long as that confidence is real, which is back to calibration.

Two details are easy to miss:

① After escalation, the large model should not get unlimited freedom. One implementation that hands uncertain rows to an LLM forces that LLM to choose from the same label set. Without that, the small model’s “unsure” gets converted into an output your pipeline cannot consume.

② Once a choice is made, hold it. One model router keeps its selection for the whole session to preserve prompt-cache continuity. Re-choosing every turn changes the prefix and throws the cache away — the money you saved on the cheaper call goes back out through the cache.

Parallel questions: what you save is input

Decision models let you put many questions into a single request; the official examples ask dozens at once.

A benchmark over 2,976 calls measured what that buys:

Item Result
8 questions in parallel vs one at a time 76–86% median saving on input tokens
Fixed overhead per request about 261 input tokens

So the saving is not “the model computes faster” — it is that the same state gets sent once. The more questions, the thinner that fixed overhead spreads. With one or two questions, the parallel/sequential difference is the same order as run-to-run noise, and not worth an architecture change.

Counterexample: batching can break ranking

This is the most counterintuitive finding on this page:

A DuckDB extension batches 40 rows by default. The benchmark found that the batched path fails its ranking-quality gate, while one row per request passes it.

The reason is reconstructible: packing 40 rows into one request asks the model to produce 40 relative orderings at once, which is harder than comparing a few candidates at a time. Cost and fidelity are in conflict here — and the cost side is visible (the bill) while the fidelity side is invisible unless you built a gate to measure it.

The general rule: parallel questions suit independent questions (classification, scoring, judging) and not questions that require items to be compared with each other.

One more precondition: the errors have to be random

Cascades work because of an implicit assumption: the small model’s mistakes are random noise, so the large model can correct them.

Independent out-of-distribution calibration testing contradicts that. Choice and Score are systematically overconfident; Boolean is systematically underconfident. A systematic bias is not noise: it makes the small model report high confidence on an entire class of inputs, so that class is never escalated and is answered wrongly every time.

The risk, then, is not “slightly lower accuracy” but one whole kind of input being consistently dropped, invisibly at the aggregate level. Finding it requires looking at accuracy by subclass, not overall.

In one sentence

A cascade saves in proportion to how much traffic sits in its high-confidence region; parallel questions save by amortising a fixed overhead. Both have a price — the first bets that the smaller model’s errors are random, the second hurts tasks that need cross-comparison.