AI Models vs Prediction Markets: The Graded Scoreboard
Updated August 13, 2026 · 12 min read · by Jake Hari
Every AI lab has a story about how well its model predicts the world. Almost nobody publishes the ledger. This page is the ledger. For months, our panel of AI models has put a price-blind probability on real Kalshi markets before they settle, every call stored before the outcome was knowable, and then graded against what actually happened. The headline finding is not flattering to the machines, which is exactly why you can trust the rest of it: on the ai models vs prediction markets question, the market is winning. Here is every seat's record anyway, because a scoreboard that only publishes when it is ahead is a press release.
The Quick Answer
Across 4,068 settled Kalshi markets, our multi-model AI panel's blend scores a Brier of 0.171 against the market's 0.164 on the same markets — the market is still the better forecaster overall, and we publish that plainly because the exceptions are where the story is. The closest seat is Grok (0.128 vs 0.122 on its 4,767 graded markets). The full seat-by-seat board, the calibration table, and named receipts are below.
All numbers as of September 27, 2026, recomputed at build time from 67,888 graded verdict rows across 4,068 settled markets · updated weekly · every verdict is logged before settlement, and nothing is ever quoted from memory.
The Leaderboard: Every Seat, Graded
Each row is one model seat. Its Brier score is computed only on the markets that seat actually priced, and the market's Brier beside it is computed on those exact same markets, so every comparison is apples to apples. Lower is better. The head-to-head column counts, market by market, whether the seat's number or the market's price landed closer to the outcome, written as wins–losses–ties.
| Seat | Graded | Seat Brier | Market Brier | Δ | Beat The Market | Win % |
|---|---|---|---|---|---|---|
| Grok | 4,767 | 0.1282 | 0.1218 | +0.0064 | 1,155–2,926–686 | 28.3% |
| Gemini Pro | 4,680 | 0.1324 | 0.1243 | +0.0081 | 896–1,540–2,244 | 36.8% |
| GPT | 5,241 | 0.1356 | 0.1230 | +0.0127 | 1,562–2,974–705 | 34.4% |
| Kimi | 2,416 | 0.1496 | 0.1338 | +0.0158 | 790–1,370–256 | 36.6% |
| Claude Fable | 1,443 | 0.1546 | 0.1382 | +0.0163 | 574–760–109 | 43.0% |
| Claude Opus | 1,486 | 0.1556 | 0.1368 | +0.0188 | 568–785–133 | 42.0% |
| GLM | 1,821 | 0.1501 | 0.1262 | +0.0239 | 516–1,071–234 | 32.5% |
| DeepSeek | 1,693 | 0.1602 | 0.1206 | +0.0395 | 551–816–326 | 40.3% |
| Claude Sonnet | 1,521 | 0.1760 | 0.1333 | +0.0427 | 566–885–70 | 39.0% |
| Gemini | 200 | 0.2783 | 0.1367 | +0.1417 | 61–130–9 | 31.9% |
Three honest notes on reading it. First, no seat beats the market on Brier over its full sample; the best seats lose narrowly, and the gap is the price of forecasting from training data against a crowd trading live information. Second, sample sizes differ because the panel's lineup has grown and changed; seats with a few hundred graded markets joined later or run on fewer boards, and small samples move around. Third, ties are markets where the seat's number and the price sat exactly the same distance from the outcome, which happens often when a model agrees with the market and rounds to the same figure.
Hottest Prediction Markets Right Now
- 2028 Democratic presidential nominee · $179M traded
- 2027 Pro Football Champion · $61M traded
- 2028 U.S. Presidential Election winner? · $59M traded
- 2028 Republican presidential nominee · $57M traded
- FedEx St. Jude Championship Winner · $54M traded
- Pro Baseball Champion · $53M traded
Every market above links to our full AI model verdict; browse them all on the OddsShopper prediction markets hub, and see how every settled call actually scored on the full graded scoreboard.
What A Brier Score Actually Is
A Brier score is the squared error of a probability forecast, averaged over every forecast made. Say 90% on things that happen and it rewards you; say 90% on things that do not and it punishes you hard. Zero is a crystal ball. A coin-flip shrug on everything scores 0.25. The market's roughly 0.115 here means Kalshi prices are, on average, very good probabilities, which matches what the accuracy research on prediction markets keeps finding. The interesting question this page tracks is not whether the market is good. It is where a model, or the blend of all of them, closes the distance.
The Record By Category
| Category | Graded | Blend Brier | Market Brier | Verdict | Best Seat (n≥20) |
|---|---|---|---|---|---|
| Sports | 1,231 | 0.1925 | 0.1847 | market ahead | DeepSeek (0.1684, n=559) |
| Politics | 102 | 0.1911 | 0.1543 | market ahead | GPT (0.1051, n=138) |
| Finance | 540 | 0.1547 | 0.1451 | market ahead | GLM (0.1006, n=220) |
| Weather | 615 | 0.1660 | 0.1602 | market ahead | Grok (0.1208, n=760) |
| Entertainment | 56 | 0.1022 | 0.0616 | market ahead | Grok (0.0340, n=238) |
| Mentions | 132 | 0.1617 | 0.1478 | market ahead | Grok (0.0299, n=35) |
| Crypto | 775 | 0.1487 | 0.1485 | even | Grok (0.1062, n=965) |
| Other | 616 | 0.1820 | 0.1745 | market ahead | Claude Fable (0.1397, n=162) |
| Culture | 1 | 0.1167 | 0.0064 | market ahead | — |
The pattern to take from this table is uniformity: the market is ahead in essentially every category, but the margin varies a lot. The blend runs closest in the catch-all and finance buckets. It sits furthest behind in politics, on a small sample, and in sports, where lineup news and late information move prices in ways a price-blind model cannot see. Where a single seat's number looks better than the blend in a category, treat it as promising rather than proven; those cells carry smaller samples, and this page re-grades them every week.
Calibration: When The Panel Says 70%, What Happens?
Brier measures error. Calibration measures honesty: when the blend said an event was 70-80% likely, how often did it actually happen?
| Panel Said | Markets | Avg Prediction | Actually Happened |
|---|---|---|---|
| 0–10% | 472 | 6.7% | 8.9% |
| 10–20% | 597 | 14.6% | 12.7% |
| 20–30% | 486 | 24.9% | 25.1% |
| 30–40% | 457 | 35.0% | 37.4% |
| 40–50% | 465 | 44.8% | 46.2% |
| 50–60% | 474 | 54.7% | 57.8% |
| 60–70% | 364 | 64.7% | 67.9% |
| 70–80% | 297 | 74.6% | 82.5% |
| 80–90% | 268 | 85.1% | 88.8% |
| 90–100% | 188 | 92.8% | 93.1% |
Two things stand out. In the tails, the panel is well calibrated: things it called near-certain happened about as often as advertised. In the 70-90% band, events happened more often than the panel said, meaning the blend has been underconfident about likely outcomes; it hedges toward 50 more than the world does. And in the 10-20% band the opposite shows up, with fewer hits than predicted. That hedging bias is a known large-language-model habit, it is measurable here on real settled markets, and it is one concrete reason the market's sharper numbers keep winning the Brier column.
The Research Findings
- The mentions family is close, and the panel still trails. On 150 graded “will he say it” contracts, the blend's Brier of 0.1464 runs nearer the market's 0.1302 than in most categories, without passing it. These markets reward reading a speaker's verbal habits, which is text analysis, which should be the one thing a language model is built for; the graded rows say the market still reads the room better.
- The always-YES stress test. 50.7% of graded mention contracts settled Yes, so a strategy that mindlessly answers Yes on every one scores a Brier of 0.4933 here. Any seat claiming mention-market skill has to beat that dumb baseline, not just the market; the table above shows which ones do.
- When the panel actually disagrees with the market (blend at least 10 points off the price), it has been right 257 times and wrong 407 times on 664 settled calls. Divergence is an editorial signal, not an automatic edge, and the misses below are part of the record.
- And the harder it disagrees, the worse it does. 38.7% at 10¢ (664 calls) → 36.7% at 15¢ (360 calls) → 36.6% at 20¢ (232 calls) → 38.0% at 30¢ (108 calls). The slope is not monotonic in this rebuild, which is itself worth watching — it was strictly falling when this was first measured.
Receipts: Calls The Panel Got Right
Every receipt below is a settled market where the blend stood at least ten points from the price and the outcome broke the panel's way. The reasoning quotes are the models' own stored rationales, written before settlement, pulled verbatim from the ledger.
- Will the Nasdaq-100 be above 30179.99 at the end of Sep 24, 2026 at 10am EDT? (settled Yes). Market price 43¢; panel blend 95%. Gemini Pro put it at 98% before settlement, reasoning: “With morning futures stable and the index approximately 270 points above the threshold, a sudden near-1% drop in the opening 30 minutes of trading is highly improbable.”
- Will the WTI crude oil settlement price be above 98.49 USD/Bbl on Sep 21, 2026? (settled No). Market price 71¢; panel blend 21%. Claude Fable put it at 9% before settlement, reasoning: “Nov contract needs a ~2.5% one-day rally against a three-session downtrend driven by Saudi pipeline restoration news; even Friday's high ($98.01) fell short of the strike.”
- Will the maximum temperature be 92-93° on Aug 9, 2026? (settled No). Market price 78¢; panel blend 3%.
- Garcia-Perez vs Ryser: Valentina Ryser wins (settled No). Market price 72¢; panel blend 1%.
The Misses, Because A Scoreboard Without Them Is An Ad
- Will the Nasdaq-100 be above 30319.99 at the end of Sep 24, 2026 at 10am EDT? (settled No). Market price 11¢; panel blend 81%. Grok put it at 90% before settlement, reasoning: “NDX sits ~150 points above the threshold with under three hours to 10am EDT settlement, making a breach unlikely.”
- WTI Oil 15 min · $90.57 target: WTI Oil price up in next 15 mins? (leg 45) (settled Yes). Market price 78¢; panel blend 9%. Kimi put it at 4% before settlement, reasoning: “With price at $90.35 four minutes into the window, below the $90.57 threshold and the session high, and falling to $89.92 afterward, a close at/above the reference was highly unlikely.”
- Will OpenAI release GPT-6 before Sep 16, 2026? (settled Yes). Market price 89¢; panel blend 23%. Grok put it at 16% before settlement, reasoning: “Sep 1 Path-to-Astra cleared Critical safeguards and said release is soon, but GPT-6 naming stays unresolved.”
The LeBron James rows are the most instructive failure in the ledger. The panel was asked while the market already sat at 99 cents, because the news had broken and the market had already repriced. The models, reasoning from a world that ended the day their training data did, coherently argued for a world that no longer existed. That is the whole moral of this page in one market: when the question is about fresh information, the market's edge is structural, and no amount of eloquent reasoning closes it. Model edges, where they exist, live in pattern-reading, not news.
Stay ahead of the markets.
Daily insights and expert picks on Kalshi, Polymarket, and what's moving markets.
Free forever. Unsubscribe anytime.
How The Grading Works
The pipeline is the same for every market, every week:
- Price-blind cards. Each model receives a data card with the market's real settlement rules and deterministically fetched facts, and never the market price. It returns a probability and a written rationale.
- A revision round. Seats then read each other's anonymized reasoning and may revise; revised numbers are stored separately, and the blend is the equal-weight mean of the seats.
- Logged before settlement. Every verdict row is written to the ledger with its timestamp and the market price at ask time, before the outcome exists. A market settles Yes or No per its own rules text, an outcome checker grades it, and this page recomputes every statistic from the raw rows on each rebuild.
- Reporting hygiene. Near-certain longshot legs where model and market both said under 3% and the answer was No are excluded from scored tables, because piling up gimme points is how a leaderboard lies. The seat table needs 20 graded markets before a seat appears.
The blend-level version of this record, with the ten biggest wins and live open divergences, lives on the model verdict scoreboard; this page is the seat-level research layer on the same ledger. The panel's current open calls with 3x-the-market upside are on Kalshi longshot picks, the markets where traders themselves are split hardest are in the most contested prediction markets, and the settlement fine print that decides several of these grades gets read weekly in Kalshi Fine Print Watch.
AI Forecasting vs Markets FAQ
Can AI beat prediction markets? Not overall, on our ledger. Every seat trails the market's Brier score on its full graded sample. The blend gets close in some categories, single seats beat the market head-to-head on a quarter to nearly half of the markets they price, and the panel's calls that diverged from the price have been right a minority of the time; the findings section above carries the current counts. That is a real research result, not a sales pitch.
Why grade against the market instead of just reporting accuracy? Because raw accuracy flatters everyone. Most markets resolve the way they were priced to resolve, so a forecaster can look brilliant by agreeing with the price. Grading against the market on the same markets isolates the only question that matters: did the model add information the price did not already have?
What would change this scoreboard's verdict? A seat sustaining a lower Brier than the market over a thousand-plus graded markets, or a category where the blend's lead survives growing samples. The weekly rebuild means either would show up here without anyone deciding to announce it.
The Bottom Line
Prediction markets are the strongest public forecasting machine anyone has built, and this ledger says so with our own models' report card. The value of running the panel anyway is the map it produces: where machines hedge, where they read patterns well, and exactly how large the market's information edge is, measured in Brier points instead of vibes. Leaderboards that launch empty, or gate their grades behind a signup form, are asking you to trust a scoreboard nobody can read. Ours is above, misses included, and it re-grades itself every week whether the numbers flatter us or not.
Event contracts are CFTC-regulated financial products offered on exchanges such as Kalshi, available to eligible U.S. residents 18+. The panel's probabilities are model estimates, not predictions of fact and not financial advice; market prices move constantly, and nothing here suggests any forecast or strategy will be profitable.



