Routing, voting, and cascading all hand back one model's answer. So your ceiling is set by how often every model is wrong at once. Call that β. It is not how often models agree (ρ).
Updated September 2026. The first version of this page reported a co-failure tail on MATH-500 (β = 0.052, 17 of 330). An audit traced every one of those questions to an evaluation error (16 grading defects, 1 truncated answer); the corrected numbers are below, and what the error looked like to ρ is section 03. Details: DATA_AUDIT.md.
Picture a panel of experts where you can only return one expert's answer. Choosing the best one helps, right up until a question lands on a blind spot they all share. Then no rule wins, because the right answer was never in the room.
That is the ceiling, and it is exact. Give a query to a pool of m models. If every one is wrong, no selection policy (router, weighted vote, cascade, debate) can be right, since each returns one member's answer. Accuracy is capped at 1−β, where β = P(all m wrong).
The field reports pairwise correlation ρ instead, and ρ is provably blind to β. You can hold the entire pairwise law fixed and still move β, a Fréchet-class fact we make exact in the paper. A single-factor copula calibrated on ρ underprices any common-mode failure, a bias that grows with pool size, because no pairwise number represents a shared blind spot. On today's frontier, once the grading is right, such blind spots turn out to be rare: the ceiling is high, and what limits combining models is knowing which one to trust.
Grade the models once on a held-out set and count the questions all of them missed. That count alone caps what any router could add. No training, no cost. Move the inputs and watch the ceiling.
Our first release saw 17 MATH-500 questions that every model got wrong (16 after the truncation control). Every one of those 16 was a grading error: a reference answer such as “x=5” or “864 inches²” that no response could match. To a pairwise statistic that looks like a shared blind spot, and the underpricing grows with the pool. Drag the slider to see it open up.
scroll to pan ↔
After the audit, no MATH-500, MATH-Hard, AIME, GPQA or competition-code question defeats every model. The ceiling is near 1. Where the best model falls short, combining could help in principle, but only by knowing per question which model to trust.
Take hard science questions and remove the options. Accuracy changes little, and the models miss different questions: on 54 well-posed questions, none stumps them all. (Our first release reported 10 of 79; most of those could not be answered without their options.)
scroll to pan ↔
Every number here recomputes live over one 2026 OpenRouter pool, from $30/Mtok flagships down to $0.03/Mtok open weights. The roster, the matrices, the grading, and the code are all released to rerun.