AI Models vs Prediction Markets: The Graded Scoreboard
Updated August 13, 2026 · 12 min read · by Jake Hari
Every AI lab has a story about how well its model predicts the world. Almost nobody publishes the ledger. This page is the ledger. For months, our panel of AI models has put a price-blind probability on real Kalshi markets before they settle, every call stored before the outcome was knowable, and then graded against what actually happened. The headline finding is not flattering to the machines, which is exactly why you can trust the rest of it: on the ai models vs prediction markets question, the market is winning. Here is every seat's record anyway, because a scoreboard that only publishes when it is ahead is a press release.
The Quick Answer
Across 2,444 settled Kalshi markets, our multi-model AI panel's blend scores a Brier of 0.127 against the market's 0.115 on the same markets — the market is still the better forecaster overall, and we publish that plainly because the exceptions are where the story is. The closest seat is Grok (0.121 vs 0.114 on its 1,870 graded markets). The full seat-by-seat board, the calibration table, and named receipts are below.
All numbers as of August 13, 2026, recomputed at build time from 23,961 graded verdict rows across 2,444 settled markets · updated weekly · every verdict is logged before settlement, and nothing is ever quoted from memory.
The Leaderboard: Every Seat, Graded
Each row is one model seat. Its Brier score is computed only on the markets that seat actually priced, and the market's Brier beside it is computed on those exact same markets, so every comparison is apples to apples. Lower is better. The head-to-head column counts, market by market, whether the seat's number or the market's price landed closer to the outcome, written as wins–losses–ties.
| Seat | Graded | Seat Brier | Market Brier | Δ | Beat The Market | Win % |
|---|---|---|---|---|---|---|
| Grok | 1,870 | 0.1214 | 0.1143 | +0.0072 | 394–1,207–269 | 24.6% |
| Gemini Pro | 1,803 | 0.1276 | 0.1154 | +0.0122 | 289–622–892 | 31.7% |
| GPT | 2,249 | 0.1320 | 0.1156 | +0.0164 | 603–1,298–348 | 31.7% |
| Kimi | 1,199 | 0.1423 | 0.1202 | +0.0221 | 291–683–225 | 29.9% |
| GLM | 1,225 | 0.1403 | 0.1161 | +0.0242 | 278–723–224 | 27.8% |
| DeepSeek | 1,214 | 0.1629 | 0.1180 | +0.0449 | 323–572–319 | 36.1% |
| Claude Fable | 235 | 0.1713 | 0.1248 | +0.0465 | 104–121–10 | 46.2% |
| Claude Opus | 235 | 0.1827 | 0.1248 | +0.0579 | 84–142–9 | 37.2% |
| Claude Sonnet | 243 | 0.1840 | 0.1207 | +0.0632 | 77–157–9 | 32.9% |
| Gemini | 184 | 0.2914 | 0.1447 | +0.1467 | 55–120–9 | 31.4% |
Three honest notes on reading it. First, no seat beats the market on Brier over its full sample; the best seats lose narrowly, and the gap is the price of forecasting from training data against a crowd trading live information. Second, sample sizes differ because the panel's lineup has grown and changed; seats with a few hundred graded markets joined later or run on fewer boards, and small samples move around. Third, ties are markets where the seat's number and the price sat exactly the same distance from the outcome, which happens often when a model agrees with the market and rounds to the same figure.
Hottest Prediction Markets Right Now
- 2028 Democratic presidential nominee · $179M traded
- 2027 Pro Football Champion · $61M traded
- 2028 U.S. Presidential Election winner? · $59M traded
- 2028 Republican presidential nominee · $57M traded
- FedEx St. Jude Championship Winner · $54M traded
- Pro Baseball Champion · $53M traded
Every market above links to our full AI model verdict; browse them all on the OddsShopper prediction markets hub, and see how every settled call actually scored on the full graded scoreboard.
What A Brier Score Actually Is
A Brier score is the squared error of a probability forecast, averaged over every forecast made. Say 90% on things that happen and it rewards you; say 90% on things that do not and it punishes you hard. Zero is a crystal ball. A coin-flip shrug on everything scores 0.25. The market's roughly 0.115 here means Kalshi prices are, on average, very good probabilities, which matches what the accuracy research on prediction markets keeps finding. The interesting question this page tracks is not whether the market is good. It is where a model, or the blend of all of them, closes the distance.
The Record By Category
| Category | Graded | Blend Brier | Market Brier | Verdict | Best Seat (n≥20) |
|---|---|---|---|---|---|
| Sports | 565 | 0.1708 | 0.1509 | market ahead | Grok (0.1675, n=347) |
| Politics | 41 | 0.1084 | 0.0836 | market ahead | Gemini Pro (0.0906, n=32) |
| Finance | 316 | 0.1215 | 0.1146 | market ahead | Kimi (0.1038, n=178) |
| Weather | 403 | 0.1132 | 0.1002 | market ahead | Grok (0.1114, n=327) |
| Entertainment | 130 | 0.0369 | 0.0255 | market ahead | Grok (0.0411, n=118) |
| Mentions | 89 | 0.1643 | 0.1452 | market ahead | Kimi (0.1833, n=88) |
| Crypto | 513 | 0.1016 | 0.0942 | market ahead | Gemini Pro (0.1003, n=406) |
| Other | 387 | 0.1390 | 0.1344 | even | GPT (0.1436, n=365) |
The pattern to take from this table is uniformity: the market is ahead in essentially every category, but the margin varies a lot. The blend runs closest in the catch-all and finance buckets. It sits furthest behind in politics, on a small sample, and in sports, where lineup news and late information move prices in ways a price-blind model cannot see. Where a single seat's number looks better than the blend in a category, treat it as promising rather than proven; those cells carry smaller samples, and this page re-grades them every week.
Calibration: When The Panel Says 70%, What Happens?
Brier measures error. Calibration measures honesty: when the blend said an event was 70-80% likely, how often did it actually happen?
| Panel Said | Markets | Avg Prediction | Actually Happened |
|---|---|---|---|
| 0–10% | 377 | 6.0% | 5.3% |
| 10–20% | 295 | 14.5% | 8.5% |
| 20–30% | 223 | 24.9% | 21.5% |
| 30–40% | 176 | 35.3% | 33.0% |
| 40–50% | 214 | 44.7% | 51.4% |
| 50–60% | 218 | 54.8% | 58.7% |
| 60–70% | 160 | 64.8% | 65.0% |
| 70–80% | 161 | 74.7% | 85.7% |
| 80–90% | 151 | 85.1% | 92.1% |
| 90–100% | 469 | 96.2% | 96.8% |
Two things stand out. In the tails, the panel is well calibrated: things it called near-certain happened about as often as advertised. In the 70-90% band, events happened more often than the panel said, meaning the blend has been underconfident about likely outcomes; it hedges toward 50 more than the world does. And in the 10-20% band the opposite shows up, with fewer hits than predicted. That hedging bias is a known large-language-model habit, it is measurable here on real settled markets, and it is one concrete reason the market's sharper numbers keep winning the Brier column.
The Research Findings
- The mentions family is close, and the panel still trails. On 89 graded “will he say it” contracts, the blend's Brier of 0.1643 runs nearer the market's 0.1452 than in most categories, without passing it. These markets reward reading a speaker's verbal habits, which is text analysis, which should be the one thing a language model is built for; the graded rows say the market still reads the room better.
- The always-YES stress test. 70.8% of graded mention contracts settled Yes, so a strategy that mindlessly answers Yes on every one scores a Brier of 0.2921 here. Any seat claiming mention-market skill has to beat that dumb baseline, not just the market; the table above shows which ones do.
- When the panel actually disagrees with the market (blend at least 10 points off the price), it has been right 85 times and wrong 293 times on 378 settled calls. Divergence is an editorial signal, not an automatic edge, and the misses below are part of the record.
Receipts: Calls The Panel Got Right
Every receipt below is a settled market where the blend stood at least ten points from the price and the outcome broke the panel's way. The reasoning quotes are the models' own stored rationales, written before settlement, pulled verbatim from the ledger.
- Will Trump say 'MAGA / Make America Great Again' during the White House Correspondents Dinner (Jul 24, 2026)? (settled Yes). Market price 39¢; panel blend 81%. DeepSeek put it at 92% before settlement, reasoning: “Trump uses this slogan in nearly every public speech; his first correspondents' dinner after an assassination attempt is a dramatic moment he'll likely punctuate with his signature phrase.”
- Will Trump say 'IQ / Genius' during the White House Correspondents Dinner (Jul 24, 2026)? (settled Yes). Market price 28¢; panel blend 64%. Claude Opus put it at 90% before settlement, reasoning: “Trump uses 'genius'/'IQ' habitually in extended remarks, and a first-ever hostile-crowd correspondents' dinner speech gives him ample unscripted runway.”
- Will the maximum temperature be 92-93° on Aug 9, 2026? (settled No). Market price 78¢; panel blend 3%.
- Will the high temp in NYC be 89-90° on Aug 7, 2026? (settled No). Market price 86¢; panel blend 36%.
The Misses, Because A Scoreboard Without Them Is An Ad
- LeBron James Next Team: Philadelphia (settled Yes). Market price 99¢; panel blend 2%. DeepSeek put it at 0% before settlement, reasoning: “LeBron at 41 with Bronny on the Lakers and no ties to Philly makes a move to a non-contending, cold-weather team extremely unlikely.”
- Barcelona vs Macara Winner? (settled Yes). Market price 90¢; panel blend 23%.
- Will the minimum temperature be 67-68° on Aug 9, 2026? (settled Yes). Market price 68¢; panel blend 2%.
The LeBron James rows are the most instructive failure in the ledger. The panel was asked while the market already sat at 99 cents, because the news had broken and the market had already repriced. The models, reasoning from a world that ended the day their training data did, coherently argued for a world that no longer existed. That is the whole moral of this page in one market: when the question is about fresh information, the market's edge is structural, and no amount of eloquent reasoning closes it. Model edges, where they exist, live in pattern-reading, not news.
Free: The Weekly PM Market Brief — this scoreboard's weekly movers, the AI panel's newest graded calls, and the biggest market-vs-model gaps. One email, Sundays. The signup box is at the bottom of this page.
How The Grading Works
The pipeline is the same for every market, every week:
- Price-blind cards. Each model receives a data card with the market's real settlement rules and deterministically fetched facts, and never the market price. It returns a probability and a written rationale.
- A revision round. Seats then read each other's anonymized reasoning and may revise; revised numbers are stored separately, and the blend is the equal-weight mean of the seats.
- Logged before settlement. Every verdict row is written to the ledger with its timestamp and the market price at ask time, before the outcome exists. A market settles Yes or No per its own rules text, an outcome checker grades it, and this page recomputes every statistic from the raw rows on each rebuild.
- Reporting hygiene. Near-certain longshot legs where model and market both said under 3% and the answer was No are excluded from scored tables, because piling up gimme points is how a leaderboard lies. The seat table needs 20 graded markets before a seat appears.
The blend-level version of this record, with the ten biggest wins and live open divergences, lives on the model verdict scoreboard; this page is the seat-level research layer on the same ledger. The panel's current open calls with 3x-the-market upside are on Kalshi longshot picks, the markets where traders themselves are split hardest are in the most contested prediction markets, and the settlement fine print that decides several of these grades gets read weekly in Kalshi Fine Print Watch.
AI Forecasting vs Markets FAQ
Can AI beat prediction markets? Not overall, on our ledger. Every seat trails the market's Brier score on its full graded sample. The blend gets close in some categories, single seats beat the market head-to-head on a quarter to nearly half of the markets they price, and the panel's calls that diverged from the price have been right a minority of the time; the findings section above carries the current counts. That is a real research result, not a sales pitch.
Why grade against the market instead of just reporting accuracy? Because raw accuracy flatters everyone. Most markets resolve the way they were priced to resolve, so a forecaster can look brilliant by agreeing with the price. Grading against the market on the same markets isolates the only question that matters: did the model add information the price did not already have?
What would change this scoreboard's verdict? A seat sustaining a lower Brier than the market over a thousand-plus graded markets, or a category where the blend's lead survives growing samples. The weekly rebuild means either would show up here without anyone deciding to announce it.
The Bottom Line
Prediction markets are the strongest public forecasting machine anyone has built, and this ledger says so with our own models' report card. The value of running the panel anyway is the map it produces: where machines hedge, where they read patterns well, and exactly how large the market's information edge is, measured in Brier points instead of vibes. Leaderboards that launch empty, or gate their grades behind a signup form, are asking you to trust a scoreboard nobody can read. Ours is above, misses included, and it re-grades itself every week whether the numbers flatter us or not.
Event contracts are CFTC-regulated financial products offered on exchanges such as Kalshi, available to eligible U.S. residents 18+. The panel's probabilities are model estimates, not predictions of fact and not financial advice; market prices move constantly, and nothing here suggests any forecast or strategy will be profitable.



