Model panel insights
What each seat of the verdict panel actually predicted, per report — computed from the published per-claim provenance (panel_by_role), not from the reconciled verdicts. Disagreement between seats is the pipeline's error-catching mechanism: an escalation means the proposer and critic differed and the arbiter decided; a Severity-Classifier override means a second-stage model re-graded a boundary call. Both are disclosed per claim on its provenance strip.
George W. Bush — 2006-01-31
48 claims · 4 escalated to the arbiter (8.3%)
| Seat | Model | Seat predictions | False-rate | Claims voted |
|---|---|---|---|---|
| Proposer | opus-worker | True 33 · Misleading 3 · False 1 · Unverifiable 1 | 2.6% | 38 |
| Critic | grok-4.3 | True 31 · Misleading 3 · False 3 · Unverifiable 1 | 7.9% | 38 |
| Arbiter | gpt-5.5 | True 1 · Misleading 2 · False 1 | 25.0% | 4 |
Arbiter side-taking on the 4 escalated claims: sided with the proposer 2, with the critic 2, took a third position 0.
Severity-Classifier stage-2 overrides: False→Misleading 1.
Bill Clinton — 1998-01-27
92 claims · 15 escalated to the arbiter (16.3%)
| Seat | Model | Seat predictions | False-rate | Claims voted |
|---|---|---|---|---|
| Proposer | opus-worker | True 66 · Misleading 5 | 0.0% | 71 |
| Critic | grok-4.3 | True 57 · Misleading 4 · False 8 · Unverifiable 2 | 11.3% | 71 |
| Arbiter | gpt-5.5 | True 8 · Misleading 4 · False 3 | 20.0% | 15 |
Arbiter side-taking on the 15 escalated claims: sided with the proposer 8, with the critic 6, took a third position 1.
Severity-Classifier stage-2 overrides: False→Misleading 2.
Joe Biden — 2022-03-01
111 claims · 11 escalated to the arbiter (9.9%)
| Seat | Model | Seat predictions | False-rate | Claims voted |
|---|---|---|---|---|
| Proposer | opus-worker | True 82 · Misleading 5 | 0.0% | 87 |
| Critic | grok-4.3 | True 74 · Misleading 4 · False 8 · Unverifiable 1 | 9.2% | 87 |
| Arbiter | gpt-5.5 | True 7 · Misleading 1 · False 2 · Unverifiable 1 | 18.2% | 11 |
Arbiter side-taking on the 11 escalated claims: sided with the proposer 6, with the critic 2, took a third position 3.
Severity-Classifier stage-2 overrides: False→Misleading 1.
Barack Obama — 2014-01-28
96 claims · 8 escalated to the arbiter (8.3%)
| Seat | Model | Seat predictions | False-rate | Claims voted |
|---|---|---|---|---|
| Proposer | opus-worker | True 71 · Misleading 5 · False 1 | 1.3% | 77 |
| Critic | grok-4.3 | True 67 · Misleading 2 · False 7 · Unverifiable 1 | 9.1% | 77 |
| Arbiter | gpt-5.5 | True 3 · Misleading 2 · False 2 · Unverifiable 1 | 25.0% | 8 |
Arbiter side-taking on the 8 escalated claims: sided with the proposer 3, with the critic 3, took a third position 2.
Severity-Classifier stage-2 overrides: False→Misleading 2.
Donald Trump — 2026-02-24
182 claims · 30 escalated to the arbiter (16.5%)
| Seat | Model | Seat predictions | False-rate | Claims voted |
|---|---|---|---|---|
| Proposer | opus-worker | True 63 · Misleading 28 · False 35 · Unverifiable 1 | 27.6% | 127 |
| Critic | grok-4.3 | True 56 · Misleading 9 · False 61 · Unverifiable 1 | 48.0% | 127 |
| Arbiter | gpt-5.5 | True 8 · Misleading 9 · False 10 · Unverifiable 3 | 33.3% | 30 |
Arbiter side-taking on the 30 escalated claims: sided with the proposer 13, with the critic 12, took a third position 5.
Severity-Classifier stage-2 overrides: Disagreement→Misleading 1 · False→Misleading 7 · Misleading→False 1.
Method notes: seat predictions are each seat's own verdict before reconciliation; the False-rate is the share of that seat's votes reading False. See About for the full pipeline.