CO · FAILURE

WHEN EVERY MODEL IS
WRONG AT ONCE, NO
ENSEMBLE CAN WIN
2026 · 06 · 15  ·  OPENROUTER FRONTIER
0.000 β = P(ALL MODELS WRONG) · MATH-500, AUDITED
0 OF 330 QUESTIONS · 67 MODELS, 21 PROVIDERS
move the cursor, the field scatters
01

If they are all wrong, no vote can win.

Routing, voting, and cascading all hand back one model's answer. So your ceiling is set by how often every model is wrong at once. Call that β. It is not how often models agree (ρ).

Updated September 2026. The first version of this page reported a co-failure tail on MATH-500 (β = 0.052, 17 of 330). An audit traced every one of those questions to an evaluation error (16 grading defects, 1 truncated answer); the corrected numbers are below, and what the error looked like to ρ is section 03. Details: DATA_AUDIT.md.

Picture a panel of experts where you can only return one expert's answer. Choosing the best one helps, right up until a question lands on a blind spot they all share. Then no rule wins, because the right answer was never in the room.

That is the ceiling, and it is exact. Give a query to a pool of m models. If every one is wrong, no selection policy (router, weighted vote, cascade, debate) can be right, since each returns one member's answer. Accuracy is capped at 1−β, where β = P(all m wrong).

The field reports pairwise correlation ρ instead, and ρ is provably blind to β. You can hold the entire pairwise law fixed and still move β, a Fréchet-class fact we make exact in the paper. A single-factor copula calibrated on ρ underprices any common-mode failure, a bias that grows with pool size, because no pairwise number represents a shared blind spot. On today's frontier, once the grading is right, such blind spots turn out to be rare: the ceiling is high, and what limits combining models is knowing which one to trust.

02

Know your ceiling before you build it.

Grade the models once on a held-out set and count the questions all of them missed. That count alone caps what any router could add. No training, no cost. Move the inputs and watch the ceiling.

FIG 1 · Realizability certificateMATH-500 default
·β̂ = K/n
·certified ceiling 1−β_lo (95% CP)
·certified max gain
Verdict pending
Defaults are the paper's audited MATH-500 run (K=0, n=330, 67 models, single best 0.988 entered at its lower confidence limit 0.952, β̂ = ·). The Clopper-Pearson lower bound on β turns the count K/n into a certified ceiling 1−β_lo on achievable accuracy, the most any router, vote, or cascade could reach. Subtracting the single-best accuracy upper-bounds the gain of every selection policy, from one labelled sample with no router trained. β_lo is the lower end of the two-sided 95% CP interval, and the default single-best value is its Bonferroni lower limit over the 67 models (the best model is picked from the same data), so the bound holds with ≥95% confidence, as in the paper; enter your own lower limit when you use it. With these defaults the certificate cannot rule out a gain of up to 0.048, the price of certifying with 330 questions and 67 candidate best models; the point estimate of the oracle gain is 0.012.
03

What a grading error looks like to ρ.

Our first release saw 17 MATH-500 questions that every model got wrong (16 after the truncation control). Every one of those 16 was a grading error: a reference answer such as “x=5” or “864 inches²” that no response could match. To a pairwise statistic that looks like a shared blind spot, and the underpricing grows with the pool. Drag the slider to see it open up.

FIG 2 · An artifact atom (June MATH-500 grading)random sub-pools
k = 67

scroll to pan ↔

The June tail (every question a grading error) over the tetrachoric single-factor prediction: median across 200 random k-model sub-pools per size, 5-95% band; at the full pool there is one pool. At k=67 the “underpricing” is ·. The signature identifies a common mode, not its cause; reading the all-wrong questions does.
04

The ceiling, benchmark by benchmark.

After the audit, no MATH-500, MATH-Hard, AIME, GPQA or competition-code question defeats every model. The ceiling is near 1. Where the best model falls short, combining could help in principle, but only by knowing per question which model to trust.

Audited β with the as-graded count. Genuine co-failures appear only on MMLU-Pro, asked for a direct answer (1 of 124; the blind raters also judged that question ambiguous). With special checkers for problems that accept several correct outputs, no competition-code problem defeats every model either.
05

Same questions, new format, still no shared blind spot.

Take hard science questions and remove the options. Accuracy changes little, and the models miss different questions: on 54 well-posed questions, none stumps them all. (Our first release reported 10 of 79; most of those could not be answered without their options.)

FIG 3 · Content-controlled format flip5-judge panel · κ 0.77-0.92
≈ 0co-failure β (same questions)
·mean accuracy (matched models)
·all-models-wrong items / 54

scroll to pan ↔

Each cell is one of the 54 GPQA-Diamond questions that survive a screen, fixed before the new data, for questions answerable without their options, content held fixed. Toggle the format: in both, no question is failed by every model (multiple choice 0, free response 0). Under the strictest judge rule (all five must agree) and the lenient rule the counts are (4 and 0 all-wrong).
06

The cast: 67 current models, 21 providers.

Every number here recomputes live over one 2026 OpenRouter pool, from $30/Mtok flagships down to $0.03/Mtok open weights. The roster, the matrices, the grading, and the code are all released to rerun.

67 models · 21 providers · priced live · temperature 0 · one model × question co-failure matrix per benchmark
Tier: frontier · mid · cheap / open-weight. Every instrument above draws on this pool, and all of it is released to replicate: the full roster with live prices, the model × question outcome matrices, the grading, and the analysis code. Every number on this page regenerates offline.