TL;DR
The panel's top of the board: Atlanta (84%), Philadelphia (17%), Washington (4%). Every verdict below was made price-blind from live fetched data, and every one gets graded against the real settlement on our scoreboard.
The crowd gives Atlanta 75 cents and Philadelphia 27: note those sum past a dollar, which is itself a small market story — while the fetched standings show the actual gap. The panel scored the race from the table, and the interesting tension is whether a five-and-a-half-game July lead deserves three-to-one conviction with two months and a deadline week remaining.
Free: The Weekly PM Market Brief — the 8-model panel's graded record, the week's biggest market-vs-model gaps, and what's spiking next. One email, Sundays. One email, Sundays — the signup box is at the bottom of this page, or just keep reading: the record it promises is on this one.
What The Models Were Given
Each model saw the complete live NL East standings (with games back) fetched from MLB StatsAPI at generation time, plus deadline-week context. All eight seats returned a verdict on every outcome in this board.
Where The Machines Split From The Money
Philadelphia: market 27¢, AI blend 17% (10 points below the crowd). Claude Sonnet: "Phillies trail Braves by 5.5 games with ~2 months left; Braves strongly favored to hold NL East."
Atlanta: market 75¢, AI blend 84% (9 points above the crowd). DeepSeek V4: "Atlanta leads by 5.5 games, has best record, and ~57 games remain, projecting >90% division win probability."
The Board
| Outcome | Kalshi | AI blend | ChatGPT (GPT-5.5) | Claude Fable | Claude Opus | Claude Sonnet | Gemini 3.1 Pro | GLM 5.2 | Kimi K3 | DeepSeek V4 |
|---|---|---|---|---|---|---|---|---|---|---|
| Atlanta | 75¢ | 84% | 84% | 80% | 78% | 90% | 78% | 82% | 85% | 92% |
| Philadelphia | 27¢ | 17% | 16% | 15% | 21% | 3% | 14% | 16% | 15% | 35% |
| Washington | 1¢ | 4% | 1.8% | 3% | 4% | 3% | 4% | 3% | 5% | 8% |
| Miami | 1¢ | 1.0% | 0.2% | 0.8% | 2% | 2% | 1.1% | 0.5% | 1.0% | 0.5% |
| New York M | 1¢ | 0.2% | 0.1% | 0.2% | 0.2% | 0.5% | 0.1% | 0.1% | 0.2% | 0.1% |
Every seat on this panel is graded against real market settlements — records to date: GPT 83% on 1,693 graded calls · Claude Fable 85% on 331 graded calls · Claude Opus 86% on 382 graded calls · Claude Sonnet 84% on 363 graded calls · Gemini 85% on 1,718 graded calls · GLM 80% on 1,874 graded calls · Kimi 84% on 1,863 graded calls · DeepSeek 80% on 1,877 graded calls. Recomputed daily; the full scoreboard is public.
Model estimates generated July 27, 2026, price-blind. These are model estimates, not predictions of fact and not financial or trading advice. Models are frequently wrong; the market price reflects real traders' money. Kalshi is a CFTC-regulated exchange; 18+, availability varies by state.
Related Verdicts
FAQ
Why do the model percentages differ from the Kalshi price?
The models never see the price. When they disagree with the crowd, one side is wrong, and we grade every verdict against real settlements on our scoreboard.
Are model verdicts betting advice?
No. Model verdicts are model estimates, not betting or financial advice. Treat them as one input among many and make your own decisions.
What data did the models see?
Each model saw the complete live NL East standings (with games back) fetched from MLB StatsAPI at generation time, plus deadline-week context.
Prices on this page are Kalshi's book. If you also trade on Polymarket, code OS4 gets new users a $20 bonus on a $10 deposit — affiliate link; terms as stated by Polymarket; 18+, availability varies by state.



