AI Prediction Market Scoreboard: Every Call Graded In Public
This page is the running, public grade sheet for every AI model verdict we publish on Kalshi prediction markets. Our panel of AI models has logged 27,089 individual model predictions across 2,686 markets, and 1,454 of those markets have settled and been graded. Every number below is recomputed from the same stored ledger, updated weekly, as of August 2, 2026. Two things up front: the models are frequently wrong, and right now the market's own prices are beating our blend on the graded set. We publish the comparison anyway, every week, because a track record you only show when you are winning is not a track record.
These are model estimates, not predictions of fact and not financial advice. Kalshi event contracts trade on a CFTC-regulated exchange; you must be 18 or older and in an eligible state to participate.
The Quick Answer
On the 1,002 graded markets where we logged a price, our blended model panel scores a Brier of 0.133 against the market's 0.116 on the exact same markets (lower is better; a permanent coin-flip scores 0.250). The market is ahead. When the blend has left the market by 10 points or more, it has been right 41 times and wrong 173 times. The full category record, the ten biggest documented wins, the five ugliest misses, and the ten biggest calls still open right now are all below.
The Honesty Block
| Model predictions logged | 27,089 individual verdicts |
| Markets covered | 2,686 distinct Kalshi markets |
| Markets graded | 1,454 settled and scored |
| Still pending | 1,232 markets awaiting settlement |
| Blend Brier (graded, priced) | 0.133 on 1,002 markets |
| Market Brier (same markets) | 0.116 |
| Plain-English verdict | the market is ahead by 0.017 Brier points |
| Numbers as of | August 2, 2026 |
Read that middle pair honestly: the crowd's money is still the best single forecaster we have measured, and any site selling you AI picks without showing you this same comparison is hiding it. The value of the panel is not that it beats the price on average. It is the specific, documented spots where it leaves the price and turns out to be right, and the public record of how often that actually happens.
The Graded Record By Category
| Category | Graded | Blend Brier | Market Brier | Who leads | Most calibrated seat |
|---|---|---|---|---|---|
| Sports | 194 | 0.200 | 0.173 | market ahead | Claude Fable (0.186, n=31) |
| Politics | 15 | 0.192 | 0.144 | market ahead | sample too small |
| Finance | 133 | 0.109 | 0.093 | market ahead | Grok (0.106, n=120) |
| Weather | 150 | 0.128 | 0.116 | market ahead | Grok (0.113, n=136) |
| Entertainment | 69 | 0.051 | 0.036 | market ahead | Grok (0.048, n=70) |
| Mentions | 89 | 0.164 | 0.145 | market ahead | Kimi (0.183, n=88) |
| Crypto | 208 | 0.113 | 0.097 | market ahead | GPT (0.109, n=183) |
| Other | 144 | 0.115 | 0.107 | market ahead | GLM (0.115, n=137) |
A seat needs at least 20 graded markets in a category to appear in the last column, and even then these leads are promising, not proven: at 20 to 90 graded markets a hot streak and real calibration look identical. We will keep grading until the difference is boring.
Receipts: The 10 Biggest Wins
A win here has a strict definition: the blend diverged from the market price by 10 points or more, we logged it before settlement, and the market settled on our side. Sorted by the size of the divergence. One disclosure: most of these divergences come from the fast-settling calibration boards we grade daily (crypto, weather, mentions), where the panel takes hundreds of small swings; the long-dated article boards settle much more slowly and are underrepresented so far.
| Market | Market price | Our blend | Settled | Divergence |
|---|---|---|---|---|
| Will July 14 be the day with the most transit calls through the Strait of Hormuz (7/13 - 7/19)? | 90c | 44% | NO | 46 pts |
| Will Trump say 'MAGA / Make America Great Again' during the White House Correspondents Dinner (Jul 24, 2026)? | 39c | 81% | YES | 42 pts |
| Will the high temp in NYC be 79-80° on Jul 20, 2026? | 77c | 37% | NO | 40 pts |
| Will the minimum temperature be >83° on Aug 1, 2026? | 51c | 12% | NO | 39 pts |
| Will Trump say 'IQ / Genius' during the White House Correspondents Dinner (Jul 24, 2026)? | 28c | 64% | YES | 36 pts |
| Will Trump say 'Kamala' during the White House Correspondents Dinner (Jul 24, 2026)? | 31c | 67% | YES | 36 pts |
| Will Trump say 'Newsom / Newscum' during the White House Correspondents Dinner (Jul 24, 2026)? | 19c | 50% | YES | 31 pts |
| Will the Nasdaq-100 be above 28959.99 at the end of Jul 22, 2026 at 4pm EDT? | 43c | 74% | YES | 31 pts |
| Will Trump say 'Israel / Israeli' during the White House Correspondents Dinner (Jul 24, 2026)? | 30c | 60% | YES | 30 pts |
| How long will Donald Trump speak for at Remarks in Marietta, Georgia? (leg 59) | 27c | 56% | YES | 29 pts |
The 5 Biggest Misses
The misses stay on the page permanently. They are the reason you can trust the wins.
| Market | Market price | Our blend | Settled | Divergence |
|---|---|---|---|---|
| LeBron James Next Team: Philadelphia | 99c | 2% | YES | 97 pts |
| LeBron James Next Team: Stays with Los Angeles L or Retires | 1c | 84% | NO | 83 pts |
| Will the maximum temperature be 95-96° on Jul 20, 2026? (leg B95.5) | 100c | 37% | YES | 63 pts |
| Will the temp in Washington DC be above 77.99° on Jul 21, 2026 at 4am EDT? | 100c | 42% | YES | 58 pts |
| Will OKSavingsBank BRION win the Gen.G vs. OKSavingsBank BRION League of Legends match? | 80c | 24% | YES | 56 pts |
The ledger's ugliest entry, in plain terms: the panel put 2% on “LeBron James Next Team: Philadelphia” while the market sat at 99c, and the market was right. That row is exactly why the “Will July 14 be the day with the most transit calls through the Strait of Hormuz (7/13 - 7/19)?” win above means something.
Live Calls: The Biggest Open Divergences Right Now
These are the largest gaps between a published model blend and the live Kalshi price at this scoreboard's last refresh (live prices as of August 2, 2026). Each one is a standing claim that will be graded onto this page when it settles. A positive gap means the blend is higher than the price.
| Market | Our blend | Live price | Gap | Settles | Full verdict |
|---|---|---|---|---|---|
| Will the Fed cut rates 0 times? | 27% | 86c | -59 pts | 2026-12-31 | our article |
| Will Shohei Ohtani win NL MVP? | 15% | 62c | -47 pts | 2026-12-31 | our article |
| Will Pete Crow-Armstrong win NL MVP? | 9% | 36c | -27 pts | 2026-12-31 | our article |
| Will Seattle be the 2026 AL West Division Winner | 23% | 34c | -11 pts | 2026-11-01 | our article |
| Will San Antonio win the Pro Basketball Western Conference Finals in the 2026-27 season? | 15% | 34c | -20 pts | 2027-06-30 | our article |
| Will the Fed cut rates 2 times? | 22% | 3c | +20 pts | 2026-12-31 | our article |
| Will House Control be Democratic AND Senate Control be Democratic for Feb 2027? | 24% | 46c | -21 pts | 2027-02-01 | our article |
| Will Miami win the Pro Basketball Eastern Conference Finals in the 2026-27 season? | 25% | 9c | +16 pts | 2027-06-30 | our article |
| Will Texas Tech go undefeated in the 2026 College Football regular season? | 8% | 27c | -19 pts | 2026-12-14 | our article |
| Will the maximum WTI front month settle price reach $150.01 by Dec 31, 2026? | 4% | 14c | -11 pts | 2027-01-01 | our article |
How The Scoreboard Works
- Price-blind protocol. Each model in the panel is asked for a probability before it is shown the market price, working from a fetched data card of verified facts (rosters, ledgers, filings, schedules). Cards are fetched, never hand-written.
- Revision rounds. Panels run structured revision rounds; the stored verdict is the one the article shipped with, timestamped at ask time, before settlement. Every prediction on this page was published before the market settled.
- Grading rule. Brier score: the squared gap between the stated probability and the outcome (YES=1, NO=0), averaged. Zero is perfect, 0.250 is a permanent coin-flip, lower is better. The market is scored on the price logged at the same moment the models were asked, on the exact same markets.
- Trivial legs are excluded. Markets where the model said 3% or less, the price was 3c or less, and the outcome was NO are dropped from every scored table. Predicting that a 1c longshot loses is not skill, and leaving those rows in would flatter every number on this page.
- The universe, pinned. Every stored panel verdict row with a probability, averaged per model per market across runs; the blend is the equal-weight mean of the model seats. Rows stay in the database forever; nothing is retroactively removed from the ledger.
- Refresh cadence. A weekly job recomputes this entire page from the ledger and republishes it only when new markets have been graded. The page you are reading was generated by that job, not edited by hand.
A Worked Example: How One Call Gets Graded
Take the top win above: “Will July 14 be the day with the most transit calls through the Strait of Hormuz (7/13 - 7/19)?” The blend said 44% while the market priced it at 90c, and it settled NO. Score both against the outcome (0): the blend takes (0.44 minus 0) squared = 0.194, the market takes (0.90 minus 0) squared = 0.810. Lower is better, so the blend won this market by 0.616 Brier points. Every graded market on this page went through exactly that arithmetic, and the headline numbers are the averages.
Every individual model verdict article carries its own market table, and each one feeds this scoreboard. If you want the methodology in action, start with the prediction markets hub or read how prediction markets work and how accurate prediction markets are. And when a sports call points you at a sportsbook number instead of an event contract, shop it first on the live odds screen.
To be explicit one last time: these are model estimates, not predictions of fact and not financial advice. The models are frequently wrong, the market is currently outscoring the blend, and this page exists so you never have to take our word for either claim. Kalshi is a CFTC-regulated exchange; event contracts are for adults 18+ and availability varies by state.
Scoreboard recomputed from the graded verdict ledger as of August 2, 2026. 18+, availability varies by state.



