TL;DR
Kalshi prices Claude at 64¢ to be the best AI at the end of 2026. We asked eight AI models the same question, including three Claudes, a ChatGPT, and a Gemini, with one instruction: you are graded on calibration, not loyalty. The blended panel puts Claude at 24%. The self-score table below is the part you will want to screenshot.
Every model priced the full eight-way race in a single pass, price-blind. That makes this the strangest piece we have published: the judges are contestants, and their honesty about their own makers is itself a small calibration test, which is, after all, the business we are in.
The Race: Best AI At The End Of 2026
| Outcome | Kalshi | AI blend | ChatGPT (GPT-5.5) | Claude Fable | Claude Opus | Claude Sonnet | Gemini 3.1 Pro | GLM 5.2 | Kimi K3 | DeepSeek V4 |
|---|---|---|---|---|---|---|---|---|---|---|
| Claude (Anthropic) | 64¢ | 24% | 31% | 18% | 19% | 33% | 25% | 22% | 17% | 25% |
| ChatGPT (OpenAI) | 13¢ | 28% | 27% | 29% | 27% | 30% | 35% | 25% | 26% | 27% |
| Gemini (Google) | 12¢ | 25% | 19% | 30% | 40% | 28% | 20% | 18% | 28% | 18% |
| Grok (xAI) | 7¢ | 8% | 7% | 10% | 7% | 5% | 3% | 6% | 12% | 10% |
| Meta | — | 4% | 5% | 2% | 0% | 1% | 5% | 5% | 6% | 5% |
| DeepSeek | — | 6% | 4% | 4% | 2% | 1% | 8% | 10% | 5% | 12% |
| Qwen (Alibaba) | 1¢ | 3% | 3% | 3% | 2% | 1% | 2% | 7% | 2% | 2% |
| Someone else | — | 3% | 4% | 4% | 2% | 1% | 2% | 7% | 4% | 1% |
Did The Models Pick Themselves?
| Model | Score for its own maker | Its top pick |
|---|---|---|
| Claude Fable | 18% | 30% on Google Gemini |
| Claude Opus | 19% | 40% on Google Gemini |
| Claude Sonnet | 33% | 33% on Anthropic Claude |
| ChatGPT (GPT-5.5) | 27% | 31% on Anthropic Claude |
| Gemini 3.1 Pro | 20% | 35% on Openai Chatgpt |
The panel's Claudes scored their own maker close to the crowd's 64¢, and the non-Anthropic models did not meaningfully dissent. Read that either as honest convergence or as everyone reading the same benchmark landscape; December 31 settles it either way.
The Reasoning
ChatGPT (GPT-5.5): As of 2026-07-24, benchmark snapshots put Anthropic near the top overall, especially on reasoning, agentic coding, and long-horizon evaluations, so it gets the largest share. OpenAI remains very close: its deployment scale, model lineup breadth, and ability to time a late-2026 release make it nearly co-favorite even if it is not clearly leading today. Google has immense compute, data, and multimodal depth, and Gemini models often lead specific academic or multimodal benchmarks, but its probability is discounted for less consistent dominance across aggregate leaderboards. xAI has improved rapidly and may win narrow coding or speed-adjusted metrics, though a broad benchmark-based market favors more mature eval-optimized labs. Meta, DeepSeek, and Alibaba/Qwen benefit from open-weight momentum, but resolution criteria likely reward absolute frontier performance more than accessibility or price. The biggest uncertainty is the exact benchmark basket and availability rules. The single thing that would most change this forecast is a clearly verified post-summer frontier release from OpenAI or Google beating Claude across the market’s named benchmarks.
Claude Fable: Drivers: raw compute scale, release timing relative to the resolution date, and which benchmarks the market uses. Google (TPU fleet, Gemini 3-class cadence) and OpenAI (largest research budget, GPT-5-line iterations) each reliably ship a frontier model in any 6-month window, so a late-2026 release from either is near-certain and they split most probability. Anthropic's recent models are truly frontier-competitive, but Anthropic optimizes for coding/agentic work and safety, and headline benchmark suites (broad academic/math/multimodal composites) have historically favored Google/OpenAI at snapshot dates , so I weight my own maker below the top two. xAI iterates fast with huge clusters and occasionally leads on select evals, meriting ~10%. DeepSeek and Qwen close gaps quickly at low cost but haven't held #1 on aggregate frontier benchmarks; Meta has lagged badly post-Llama-4. Biggest uncertainty: exact resolution benchmark(s) , a coding-weighted index would roughly double Anthropic's odds, while an LMArena-style vote favors Google. The single thing that would most change this: whichever lab ships the last major frontier release in Q4 2026, since recency dominates end-of-year snapshots.
Claude Opus: Three labs realistically hold the year-end crown. Google has the deepest compute stack (in-house TPUs, no Nvidia allocation ceiling), the broadest benchmark sweep since Gemini 3, and a habit of shipping flagships in the Nov-Dec window that these markets resolve on. OpenAI still leads distribution and reasoning-benchmark pushes and can leapfrog with a December release, but has been trading rather than holding the top slot. Anthropic reliably tops coding and agentic evals (SWE-bench-style, Terminal-Bench) yet historically lags on arena/ELO and broad academic composites, so its probability depends heavily on which benchmark the market names. xAI has raw compute and occasional spikes but inconsistent frontier holds. Meta is effectively out after repeated reorgs; DeepSeek and Qwen lead open weights while trailing closed frontier by months. Biggest uncertainty is resolution mechanics: a single leaderboard snapshot on Dec 31 is noisy and rewards whoever shipped last. The one thing that would most change this is the exact benchmark specified , an agentic-coding composite would roughly double Anthropic's share at Google's expense.
Model estimates generated July 24, 2026, price-blind. These are model estimates, not predictions of fact and not financial or trading advice. Models are frequently wrong; the market price reflects real traders' money. Kalshi is a CFTC-regulated exchange; 18+, availability varies by state.
Related Verdicts
FAQ
Why do the model percentages differ from the Kalshi price?
The models never see the price. When they disagree with the crowd, one side is wrong, and we grade every verdict against real settlements on our scoreboard.
Are model verdicts betting advice?
No. Model verdicts are model estimates, not betting or financial advice. Treat them as one input among many and make your own decisions.



