TL;DR
Last night settled everything we published about the White House Correspondents' Dinner. On the 30 graded mention markets, the Kalshi crowd narrowly beat our seven-model panel: crowd Brier 0.238, best model (ChatGPT (GPT-5.5)) 0.240, essentially a tie, and everyone else behind. Both the panel AND the crowd badly underpriced how long he would speak. We are publishing the full scorecard anyway, because a scoreboard you only show after wins is not a scoreboard.
Yesterday morning we published price-blind AI verdicts on all 34 mention markets, the speech-duration ladder, and an afternoon gap report. By midnight, Kalshi had settled nearly all of it. This is the part most content skips: the grades.
The Headline Numbers
| Forecaster | Markets | Brier score (lower = better) | vs the market |
|---|---|---|---|
| The Kalshi crowd | 30 | 0.238 | — |
| ChatGPT (GPT-5.5) | 30 | 0.240 | tied it |
| Kimi K3 | 30 | 0.285 | lost by 0.047 |
| Claude Opus | 29 | 0.292 | lost by 0.046 |
| GLM 5.2 | 30 | 0.293 | lost by 0.055 |
| Claude Sonnet | 29 | 0.326 | lost by 0.080 |
| Claude Fable | 30 | 0.333 | lost by 0.095 |
| DeepSeek V4 | 30 | 0.369 | lost by 0.131 |
Brier score = average squared error between a stated probability and what actually happened; 0 is perfect, 0.25 is coin-flip territory. All model verdicts were made price-blind before the event.
Round one to the money. ChatGPT (GPT-5.5) effectively matched the crowd; every other seat trailed it. The mention markets settled YES 50% of the time against an average price of 44¢, so the crowd ran slightly cheap on "yes" overall, but not by enough for the panel's aggressive calls to cash.
The Biggest Calls, Graded
Sorted by how far the panel stood from the crowd at publish time:
| Phrase | Kalshi price | AI blend | Result | Closer |
|---|---|---|---|---|
| 'Event does not qualify' during the White House Correspondents Dinner (Jul 24, 2026) | 5¢ | 61% | NO | market |
| 'MAGA / Make America Great Again' during the White House Correspondents Dinner (Jul 24, 2026) | 39¢ | 81% | YES | panel |
| 'IQ / Genius' during the White House Correspondents Dinner (Jul 24, 2026) | 28¢ | 64% | YES | panel |
| 'Kamala' during the White House Correspondents Dinner (Jul 24, 2026) | 31¢ | 67% | YES | panel |
| 'Hoax' during the White House Correspondents Dinner (Jul 24, 2026) | 28¢ | 60% | NO | market |
| 'Newsom / Newscum' during the White House Correspondents Dinner (Jul 24, 2026) | 19¢ | 50% | YES | panel |
| 'Israel / Israeli' during the White House Correspondents Dinner (Jul 24, 2026) | 30¢ | 60% | YES | panel |
| 'Hottest' during the White House Correspondents Dinner (Jul 24, 2026) | 65¢ | 36% | YES | market |
| 'China' during the White House Correspondents Dinner (Jul 24, 2026) | 63¢ | 91% | YES | panel |
| 'Iran (3+ times)' during the White House Correspondents Dinner (Jul 24, 2026) | 56¢ | 30% | YES | market |
| 'Barack Hussein Obama' during the White House Correspondents Dinner (Jul 24, 2026) | 55¢ | 29% | YES | market |
| 'Afford / Affordable / Affordability' during the White House Correspondents Dinner (Jul 24, 2026) | 37¢ | 62% | NO | market |
The Duration Miss Everyone Shared
The most expensive number of the night was not a mention. The 40-plus-minute rung of the duration ladder traded at 19¢ when we published, drifted to 12¢, and settled YES — he spoke for over an hour. Our panel does not get to gloat: its blended estimate was 18%, agreeing with the crowd almost exactly. The only seat leaning the right way was Claude Fable at 33%, and even that was nowhere near the truth.
The lesson is going straight into the engine. Speech-length questions about this particular speaker have a strong historical base rate that both the market and the models ignored in favor of format-based reasoning ("dinner remarks run 20-30 minutes"). That is exactly the kind of per-category lesson our grading loop exists to encode: the next duration ladder our models see will carry the base-rate card this one lacked.
Why We Publish The Losses
The product we are building is a track record. Every verdict, win or lose, lands on the same public scoreboard and gets graded against real settlements, and 30 more graded markets joined it overnight. One event proves nothing on its own; a thousand graded verdicts will. If the models cannot beat the crowd over time, the scoreboard will say so, and that honesty is worth more than any single night's bragging rights.
Model estimates graded July 25, 2026. These are model estimates, not predictions of fact and not financial or trading advice. Models are frequently wrong; the market price reflects real traders' money. Kalshi is a CFTC-regulated exchange; 18+, availability varies by state.
FAQ
Who won, the AI models or the market?
The market, narrowly. The crowd's Brier of 0.238 edged every model; ChatGPT (GPT-5.5) effectively tied it.
Will you keep publishing graded results even when the models lose?
Yes. The scoreboard only means something if the losses are on it.
Are model verdicts betting advice?
No. Model verdicts are model estimates, not betting or financial advice. Treat them as one input among many.



