AI Models vs Prediction Markets: The Graded Scoreboard
Updated August 13, 2026 · 12 min read · by Jake Hari
Every AI lab has a story about how well its model predicts the world. Almost nobody publishes the ledger. This page is the ledger. For months, our panel of AI models has put a price-blind probability on real Kalshi markets before they settle, every call stored before the outcome was knowable, and then graded against what actually happened. The headline finding is not flattering to the machines, which is exactly why you can trust the rest of it: on the ai models vs prediction markets question, the market is winning. Here is every seat's record anyway, because a scoreboard that only publishes when it is ahead is a press release.
The Quick Answer
Across 1,908 settled Kalshi markets, our multi-model AI panel's blend scores a Brier of 0.167 against the market's 0.157 on the same markets — the market is still the better forecaster overall, and we publish that plainly because the exceptions are where the story is. The closest seat is Grok (0.118 vs 0.111 on its 2,130 graded markets). The full seat-by-seat board, the calibration table, and named receipts are below.
All numbers as of August 25, 2026, recomputed at build time from 26,458 graded verdict rows across 1,908 settled markets · updated weekly · every verdict is logged before settlement, and nothing is ever quoted from memory.
The Leaderboard: Every Seat, Graded
Each row is one model seat. Its Brier score is computed only on the markets that seat actually priced, and the market's Brier beside it is computed on those exact same markets, so every comparison is apples to apples. Lower is better. The head-to-head column counts, market by market, whether the seat's number or the market's price landed closer to the outcome, written as wins–losses–ties.
| Seat | Graded | Seat Brier | Market Brier | Δ | Beat The Market | Win % |
|---|---|---|---|---|---|---|
| Grok | 2,130 | 0.1178 | 0.1109 | +0.0068 | 441–1,370–319 | 24.4% |
| Gemini Pro | 2,059 | 0.1249 | 0.1138 | +0.0110 | 330–712–1,017 | 31.7% |
| GPT | 2,532 | 0.1303 | 0.1142 | +0.0161 | 673–1,476–383 | 31.3% |
| Kimi | 1,255 | 0.1438 | 0.1220 | +0.0219 | 308–719–228 | 30.0% |
| GLM | 1,280 | 0.1420 | 0.1181 | +0.0238 | 298–756–226 | 28.3% |
| Claude Fable | 308 | 0.1616 | 0.1266 | +0.0350 | 127–168–13 | 43.1% |
| DeepSeek | 1,268 | 0.1635 | 0.1201 | +0.0434 | 349–599–320 | 36.8% |
| Claude Opus | 308 | 0.1709 | 0.1266 | +0.0442 | 108–180–20 | 37.5% |
| Claude Sonnet | 318 | 0.1832 | 0.1227 | +0.0606 | 97–206–15 | 32.0% |
| Gemini | 188 | 0.2903 | 0.1419 | +0.1484 | 56–123–9 | 31.3% |
Three honest notes on reading it. First, no seat beats the market on Brier over its full sample; the best seats lose narrowly, and the gap is the price of forecasting from training data against a crowd trading live information. Second, sample sizes differ because the panel's lineup has grown and changed; seats with a few hundred graded markets joined later or run on fewer boards, and small samples move around. Third, ties are markets where the seat's number and the price sat exactly the same distance from the outcome, which happens often when a model agrees with the market and rounds to the same figure.
Hottest Prediction Markets Right Now
- 2028 Democratic presidential nominee · $179M traded
- 2027 Pro Football Champion · $61M traded
- 2028 U.S. Presidential Election winner? · $59M traded
- 2028 Republican presidential nominee · $57M traded
- FedEx St. Jude Championship Winner · $54M traded
- Pro Baseball Champion · $53M traded
Every market above links to our full AI model verdict; browse them all on the OddsShopper prediction markets hub, and see how every settled call actually scored on the full graded scoreboard.
What A Brier Score Actually Is
A Brier score is the squared error of a probability forecast, averaged over every forecast made. Say 90% on things that happen and it rewards you; say 90% on things that do not and it punishes you hard. Zero is a crystal ball. A coin-flip shrug on everything scores 0.25. The market's roughly 0.115 here means Kalshi prices are, on average, very good probabilities, which matches what the accuracy research on prediction markets keeps finding. The interesting question this page tracks is not whether the market is good. It is where a model, or the blend of all of them, closes the distance.
The Record By Category
| Category | Graded | Blend Brier | Market Brier | Verdict | Best Seat (n≥20) |
|---|---|---|---|---|---|
| Sports | 496 | 0.1974 | 0.1806 | market ahead | Grok (0.1673, n=370) |
| Politics | 25 | 0.1950 | 0.1606 | market ahead | Gemini Pro (0.0892, n=39) |
| Finance | 269 | 0.1570 | 0.1487 | market ahead | GLM (0.1037, n=198) |
| Weather | 318 | 0.1535 | 0.1422 | market ahead | Grok (0.1128, n=366) |
| Entertainment | 40 | 0.0850 | 0.0639 | market ahead | Grok (0.0405, n=150) |
| Mentions | 83 | 0.1713 | 0.1557 | market ahead | Kimi (0.1833, n=88) |
| Crypto | 379 | 0.1495 | 0.1460 | even | Grok (0.0956, n=493) |
| Other | 298 | 0.1696 | 0.1666 | even | Claude Opus (0.0637, n=22) |
The pattern to take from this table is uniformity: the market is ahead in essentially every category, but the margin varies a lot. The blend runs closest in the catch-all and finance buckets. It sits furthest behind in politics, on a small sample, and in sports, where lineup news and late information move prices in ways a price-blind model cannot see. Where a single seat's number looks better than the blend in a category, treat it as promising rather than proven; those cells carry smaller samples, and this page re-grades them every week.
Calibration: When The Panel Says 70%, What Happens?
Brier measures error. Calibration measures honesty: when the blend said an event was 70-80% likely, how often did it actually happen?
| Panel Said | Markets | Avg Prediction | Actually Happened |
|---|---|---|---|
| 0–10% | 195 | 7.2% | 9.7% |
| 10–20% | 296 | 14.6% | 10.5% |
| 20–30% | 232 | 25.0% | 22.4% |
| 30–40% | 188 | 35.3% | 34.6% |
| 40–50% | 230 | 44.8% | 50.4% |
| 50–60% | 222 | 54.8% | 59.0% |
| 60–70% | 168 | 64.7% | 64.3% |
| 70–80% | 158 | 74.7% | 84.2% |
| 80–90% | 130 | 85.1% | 91.5% |
| 90–100% | 89 | 92.1% | 91.0% |
Two things stand out. In the tails, the panel is well calibrated: things it called near-certain happened about as often as advertised. In the 70-90% band, events happened more often than the panel said, meaning the blend has been underconfident about likely outcomes; it hedges toward 50 more than the world does. And in the 10-20% band the opposite shows up, with fewer hits than predicted. That hedging bias is a known large-language-model habit, it is measurable here on real settled markets, and it is one concrete reason the market's sharper numbers keep winning the Brier column.
The Research Findings
- The mentions family is close, and the panel still trails. On 89 graded “will he say it” contracts, the blend's Brier of 0.1643 runs nearer the market's 0.1452 than in most categories, without passing it. These markets reward reading a speaker's verbal habits, which is text analysis, which should be the one thing a language model is built for; the graded rows say the market still reads the room better.
- The always-YES stress test. 70.8% of graded mention contracts settled Yes, so a strategy that mindlessly answers Yes on every one scores a Brier of 0.2921 here. Any seat claiming mention-market skill has to beat that dumb baseline, not just the market; the table above shows which ones do.
- When the panel actually disagrees with the market (blend at least 10 points off the price), it has been right 89 times and wrong 205 times on 294 settled calls. Divergence is an editorial signal, not an automatic edge, and the misses below are part of the record.
- And the harder it disagrees, the worse it does. 30.3% at 10¢ (294 calls) → 25.2% at 15¢ (163 calls) → 24.5% at 20¢ (110 calls) → 27.7% at 30¢ (47 calls). The slope is not monotonic in this rebuild, which is itself worth watching — it was strictly falling when this was first measured.
Receipts: Calls The Panel Got Right
Every receipt below is a settled market where the blend stood at least ten points from the price and the outcome broke the panel's way. The reasoning quotes are the models' own stored rationales, written before settlement, pulled verbatim from the ledger.
- Will Trump say 'MAGA / Make America Great Again' during the White House Correspondents Dinner (Jul 24, 2026)? (settled Yes). Market price 39¢; panel blend 81%. DeepSeek put it at 92% before settlement, reasoning: “Trump uses this slogan in nearly every public speech; his first correspondents' dinner after an assassination attempt is a dramatic moment he'll likely punctuate with his signature phrase.”
- Will Trump say 'IQ / Genius' during the White House Correspondents Dinner (Jul 24, 2026)? (settled Yes). Market price 28¢; panel blend 64%. Claude Opus put it at 90% before settlement, reasoning: “Trump uses 'genius'/'IQ' habitually in extended remarks, and a first-ever hostile-crowd correspondents' dinner speech gives him ample unscripted runway.”
- Will the maximum temperature be 92-93° on Aug 9, 2026? (settled No). Market price 78¢; panel blend 3%.
- Will the high temp in NYC be 89-90° on Aug 7, 2026? (settled No). Market price 86¢; panel blend 36%.
The Misses, Because A Scoreboard Without Them Is An Ad
- Barcelona vs Macara Winner? (settled Yes). Market price 90¢; panel blend 23%.
- Will the minimum temperature be 67-68° on Aug 9, 2026? (settled Yes). Market price 68¢; panel blend 2%.
- Palmeiras vs Internacional: Total Goals: Will over 1.5 goals be scored? (leg 2) (settled No). Market price 6¢; panel blend 68%.
The LeBron James rows are the most instructive failure in the ledger. The panel was asked while the market already sat at 99 cents, because the news had broken and the market had already repriced. The models, reasoning from a world that ended the day their training data did, coherently argued for a world that no longer existed. That is the whole moral of this page in one market: when the question is about fresh information, the market's edge is structural, and no amount of eloquent reasoning closes it. Model edges, where they exist, live in pattern-reading, not news.
Stay ahead of the markets.
Daily insights and expert picks on Kalshi, Polymarket, and what's moving markets.
Free forever. Unsubscribe anytime.
How The Grading Works
The pipeline is the same for every market, every week:
- Price-blind cards. Each model receives a data card with the market's real settlement rules and deterministically fetched facts, and never the market price. It returns a probability and a written rationale.
- A revision round. Seats then read each other's anonymized reasoning and may revise; revised numbers are stored separately, and the blend is the equal-weight mean of the seats.
- Logged before settlement. Every verdict row is written to the ledger with its timestamp and the market price at ask time, before the outcome exists. A market settles Yes or No per its own rules text, an outcome checker grades it, and this page recomputes every statistic from the raw rows on each rebuild.
- Reporting hygiene. Near-certain longshot legs where model and market both said under 3% and the answer was No are excluded from scored tables, because piling up gimme points is how a leaderboard lies. The seat table needs 20 graded markets before a seat appears.
The blend-level version of this record, with the ten biggest wins and live open divergences, lives on the model verdict scoreboard; this page is the seat-level research layer on the same ledger. The panel's current open calls with 3x-the-market upside are on Kalshi longshot picks, the markets where traders themselves are split hardest are in the most contested prediction markets, and the settlement fine print that decides several of these grades gets read weekly in Kalshi Fine Print Watch.
More on this: AI Prediction Market Scoreboard: 3,625 Kalshi Calls Graded In Public · Barcelona Vs Feyenoord Odds: Does The Favorite Survive The Fee? · Where Does Bitcoin End 2026? Kalshi's Price-Band Odds · The Longest Shot On This Board Might Be The Best-Kept Secret · 2026 Governor Races: Every Kalshi Board We Track
The Bottom Line
Prediction markets are the strongest public forecasting machine anyone has built, and this ledger says so with our own models' report card. The value of running the panel anyway is the map it produces: where machines hedge, where they read patterns well, and exactly how large the market's information edge is, measured in Brier points instead of vibes. Leaderboards that launch empty, or gate their grades behind a signup form, are asking you to trust a scoreboard nobody can read. Ours is above, misses included, and it re-grades itself every week whether the numbers flatter us or not.
Event contracts are CFTC-regulated financial products offered on exchanges such as Kalshi, available to eligible U.S. residents 18+.



