Updated August 26, 2026 · 14 min read · by Sam Smith. Panel verdicts were generated August 26, price-blind, from the LiveBench.ai settlement view and the contract terms as published by the exchange. Kalshi prices fetched August 26, 2026, 15:53 UTC.
Kalshi's science and technology section carries a contract that ignores almost everything you have ever read about the AI race. It does not care about revenue, users, funding rounds or which assistant you actually enjoy talking to. It settles on one column of one website: whether a company holds the top "Coding Average" score on LiveBench.ai at 10:00 a.m. Eastern on December 31, 2026.
Nine companies have a tradeable contract. Traders have made Anthropic a heavy favorite, and on the leaderboard that decides the whole thing, Anthropic is in fact number one today.
So we handed the board to eight AI models. Each one got the live leaderboard, the full binding contract terms and a coverage note telling it exactly what it was and was not being shown. None of them saw a market price, which is the point: a model shown the price tends to repeat the price back to you. Then each model read the other seven arguments anonymously and got one chance to change its mind.
They came back with the favorite in second place, a company near the bottom of the board sitting fourth on their own list, and a tenth outcome the board does not even sell.
Free: The Weekly PM Market Brief — the 8-model panel's graded record, the week's biggest market-vs-model gaps, and what's spiking next. One email, Sundays.
New to event markets? Our plain explanation of how prediction markets work covers the mechanics in about five minutes.
The One Column That Decides Everything
LiveBench is a benchmark that rewrites its own questions periodically so models cannot simply memorize the answers. It scores everything it evaluates across seven categories and publishes each category as its own column.
Kalshi's contract terms name exactly one of them. "The Underlying for this Contract is models on LiveBench.ai ranked by Coding Average score," the document reads, and the payout goes to whichever company "has the highest ranked model by Coding Average on [the date] at 10:00 AM ET."
Three details in that sentence do most of the work on this page.
Coding Average is only two tasks. LiveBench builds that column out of code generation and code completion. It publishes a completely separate column called Agentic Coding, built from three different tasks, and the two columns rank the same models very differently. The newest Anthropic model on the table by LiveBench's own release dates, Claude 5 Opus, is the single best model listed on Agentic Coding at 65.20. On the column that actually settles this contract it sits tenth at 81.45, more than four points behind its own stablemate. A lab can be building the best coding assistant in the world and still lose this market.
The snapshot is a moment, not a season. A model that leads for eleven months and gets passed on December 20 pays nothing, and the terms add that revisions made to the leaderboard after expiration "will not be accounted for."
A tie splits the dollar. If two listed companies are level at the top, the terms pay each side's Yes holders one dollar divided by the number tied, rounded down to the whole cent. On a table where second through sixth place are separated by under two points, that is a live outcome rather than a footnote.
Anthropic Leads By Two Points, And That Is The Whole Cushion
On the settlement view as of August 25, the top of the Coding Average column reads: Claude Fable 5 at 85.99, then GPT-5.6 Sol at 83.94, then GPT-5.2 Codex at 83.62, then GPT-5.6 Luna at 82.91.
Anthropic's lead is 2.05 points. To see how thin that is, look further down the same column: second place through sixth are separated by 1.79 points in total. Anthropic holds seven of the top fourteen slots, so the depth behind the leader is real, but the crown itself rests on one model and a gap no wider than the traffic immediately below it.
There is a wrinkle in the leader's own recent history. Anthropic announced Fable 5 on June 9 as its most capable model for ambitious coding projects. Three days after that, by the company's own account, US export controls forced it to suspend access to the model entirely, and the company said it was putting the model back online on July 1 with a new safety classifier that, in its words, "comes at the cost of flagging benign requests more often during routine coding and debugging tasks." Whether that shows up in a benchmark score is genuinely unsettled, and it split our panel harder than anything else on the board.
OpenAI Has Four Of The Top Six, And No Codex In This Generation
OpenAI does not hold the crown, but it owns more of the neighborhood than any other company: four of the top six scores, including two of the three GPT-5.6 variants released July 9.
The historical pattern is the part that moved our panel. LiveBench has published eleven question sets since June 2024, and each one is its own scored table. Across all eleven, the Coding Average lead has gone to Anthropic five times, to OpenAI five times, and to Google once. Nobody else, ever.
Two Codex-branded models account for OpenAI's three most recent turns at the top of this column. GPT-5.1 Codex Max led two consecutive question sets, then GPT-5.2 Codex led the one after that, and GPT-5.2 Codex is still sitting third overall today at 83.62 while posting only 49.39 on Agentic Coding. That profile describes a model tuned narrowly for the exact two tasks this contract settles on. There is no Codex-branded model in the 5.6 generation on the table yet, and seven of our eight panelists named that absence as the single most likely thing to flip the market before New Year's Eve.
Dark Horses The Panel Won't Dismiss: Google, Moonshot AI And The Unlisted Field
Google trades near the bottom of the board and our panel does not agree with that at all.
The case for the price is easy to see on the leaderboard. Google's best entry is Gemini 3.7 Flash at 78.89, sixteenth overall and 7.10 points off the lead, and its highest-scoring Pro model, Gemini 3.1 Pro Preview, sits twenty-eighth at 76.45. The reason is documented: TechCrunch reported on July 21 that Google DeepMind shipped three new Gemini models that day but no update to Gemini Pro, which had last been updated in February, and that Bloomberg had reported Google "was facing internal delays in launching the 3.5 Pro as it struggled to meet internal performance goals." The same report carried Google DeepMind's statement that it had begun its most ambitious pre-training run yet for Gemini 4.
That combination is why the panel keeps Google alive. Google is the only company besides the top two that has ever held this column, and one top-tier release, the kind of model a lab puts at the head of its lineup, landing in time for LiveBench to score it is a real path. The panel simply treats the window as tight.
Moonshot AI is the other name the models like more than the market does. Kimi K3 sits ninth at 81.45, only 4.54 points back, closer to the crown than anything Google, Alibaba, DeepSeek or xAI has on the table today.
Then there is the outcome the board cannot sell you. Kalshi's nine contracts are mutually exclusive with each other, but the leaderboard is not restricted to them. Fifth place on the Coding Average column belongs to Smaug-Agentic at 82.47, which LiveBench lists under Abacus.AI, a company with no contract here. It is a fine-tune of Moonshot's Kimi K3, and LiveBench credits the organization that did the tuning rather than the lab that built the base model, which is exactly the attribution the payout criterion depends on. If an unlisted organization holds the top score at the snapshot, all nine contracts pay zero. Our panel gives that 9.5%.
What Eight Models Said About The Whole Board
Here is the board, as of August 26, 2026. Market prices are the bid and the ask from the live order book, fetched at 15:53 UTC. The panel column is the seat median after the revision round. With eight seats that is the midpoint of the two middle answers, which is what keeps one outlier from dragging a blended average around.

| Company | Best model on the settlement column | Market (bid / ask) | Panel |
|---|---|---|---|
| Anthropic | Claude Fable 5, 1st | 57¢ / 60¢ | 35.0% |
| OpenAI | GPT-5.6 Sol, 2nd | 26¢ / 29¢ | 36.7% |
| xAI | Grok 4.6, 27th | 10¢ / 11¢ | 1.1% |
| Z.ai | GLM-5.2, 13th | 2¢ / 3¢ | 2.0% |
| Gemini 3.7 Flash, 16th | 1¢ / 2¢ | 6.8% | |
| Moonshot AI | Kimi K3, 9th | 1¢ / 2¢ | 5.0% |
| DeepSeek | DeepSeek V4 Pro, 26th | 0¢ / 1¢ | 2.5% |
| Alibaba | Qwen 3.6 Plus, 21st | 0¢ / 1¢ | 1.8% |
| Baidu | none on the leaderboard | 0¢ / 1¢ | 0.2% |
| No Listed Company | an unlisted organization wins | not tradeable | 9.5% |
Six of the nine contracts have a real quote on both sides. DeepSeek, Alibaba and Baidu show no bid at all, so their one-cent ask is a displayed price rather than a two-sided market with someone waiting on the other end. The nine bids sum to 97 and the nine asks to 110, which is normal spread arithmetic on a board this wide and leaves little visible room in the prices for the tenth row.
Two rows deserve narration. The first is xAI, which is the sharpest disagreement on the page in relative terms: the market has it third at ten cents bid while the panel has it eighth. xAI's best model on the settlement column is Grok 4.6 at 76.78, twenty-seventh overall and 9.21 points behind, and xAI has never led a LiveBench question set. The second is Baidu, which is not merely unranked. It does not appear anywhere in LiveBench's model registry, across the current table and every past one, so the contract asks about a company the settlement source has never scored.
Every seat's number, before and after the revision round:
| Seat | Anthropic | OpenAI | Moonshot | No listed co. | |
|---|---|---|---|---|---|
| Claude Fable 5 | 40.0 → 36.0 | 36.0 → 33.0 | 8.2 → 6.5 | 3.5 → 5.0 | 5.0 → 9.0 |
| Claude Opus 5 | 40.0 → 39.5 | 35.0 → 36.3 | 7.0 → 6.5 | 4.0 → 4.0 | 7.5 → 7.0 |
| Claude Sonnet 5 | 37.0 → 35.0 | 35.0 → 37.0 | 5.0 → 6.0 | 6.0 → 4.5 | 6.5 → 9.0 |
| ChatGPT | 36.0 → 36.5 | 32.0 → 35.0 | 9.0 → 7.0 | 6.0 → 5.0 | 6.5 → 9.0 |
| Gemini | 34.0 → 34.0 | 44.0 → 38.0 | 2.0 → 5.0 | 3.0 → 6.0 | 8.0 → 10.0 |
| GLM | 28.0 → 33.0 | 32.0 → 36.0 | 4.0 → 7.0 | 8.0 → 5.0 | 17.5 → 12.0 |
| Kimi | 36.0 → 33.0 | 32.0 → 37.0 | 7.0 → 7.0 | 5.0 → 5.0 | 11.5 → 10.0 |
| DeepSeek | 25.0 → 35.0 | 35.0 → 38.0 | 8.0 → 7.0 | 3.0 → 5.0 | 25.0 → 10.0 |
| Panel Middle | 36.0 → 35.0 | 35.0 → 36.7 | 7.0 → 6.8 | 4.5 → 5.0 | 7.7 → 9.5 |
Seat numbers are each model's own percentages as submitted. One first-round board came back summing to 100.2 rather than 100, so every seat is normalized to sum to 100 before the middle is taken, which is why the bottom row is not always the plain midpoint of the column above it.
Every seat on this panel is graded against real market settlements — records to date: Gemini 86% on 4,952 graded calls · GLM 81% on 2,544 graded calls · Kimi 83% on 2,534 graded calls · DeepSeek 80% on 2,583 graded calls. Recomputed daily; the full scoreboard is public.
These are model estimates, not predictions of fact and not financial advice. Kalshi is a CFTC-regulated exchange for event contracts, 18+ only, and availability varies. Every number in this piece gets graded in public once the market settles: the running record lives on the full graded scoreboard.
One piece of context belongs right next to those columns. When this panel disagrees with a market price by ten cents or more, the market has been right roughly two thirds of the time in our own graded history. The panel is a second opinion with a public track record attached, and on a board like this one the gap is the story rather than the recommendation.
More live boards from the same panel: the broader best AI at the end of 2026 board, where Claude trades at 69¢ on a market that is not restricted to coding · the top-ranked AI model board, where OpenAI leads at 26¢ and the settlement leaderboard is a different one entirely · and Kalshi's daily prediction-market hub for what the panel priced this morning. Prices fetched August 26, 2026.
Hottest Prediction Markets Right Now
- 2028 Democratic presidential nominee · $194M traded
- 2027 Pro Football Champion · $72M traded
- 2028 U.S. Presidential Election winner? · $62M traded
- Pro Baseball Champion · $60M traded
- 2028 Republican presidential nominee · $60M traded
- Bitcoin price at the end of 2026 · $33M traded
Every market above links to our full AI model verdict; browse them all on the OddsShopper prediction markets hub, and see how every settled call actually scored on the full graded scoreboard.
Where The Panel Changed Its Mind
The revision round moved real numbers on this board, and no money changed hands doing it. The clearest example came from the seat that started furthest from everyone else. Where the seats say FIELD below, they mean the last row of the board: nobody with a contract holds the top score, so every one of the nine pays zero.
"The FIELD deserves high probability (25%) because unlisted organizations like Abacus.AI (82.47) are highly competitive, and the board excludes major players like Meta." — DeepSeek, first round
After reading the other seven arguments, it cut that to 10 and rewrote the reasoning around a narrower mechanism, while moving Anthropic up ten points:
"FIELD deserves 10% because Abacus.AI's fine-tune already ranks fifth, demonstrating unlisted organizations can compete for this narrow column." — DeepSeek, second round
The most interesting revision was an argument nobody else had made. Four of the eight seats treated Anthropic's July 1 safety-classifier redeployment as a live threat to the leader's score. One seat went back and checked the dates:
"A, C, D and H all treated Anthropic's July 1 classifier redeploy as a live downside on the 85.99. But the timeline in the card forecloses that: Fable 5 was suspended June 12, the question set is dated June 25, and site data was refreshed August 25, so 85.99 is almost certainly the post-classifier score already." — Claude Opus 5, second round
[Editor's note: LiveBench does not publish a run date beside any individual model score, so this cannot be confirmed from the settlement source. The dates do lean the way the quote does. The question set is dated June 25, Anthropic says access was restored July 1, and the site refreshed its data on August 25, which leaves most of the plausible evaluation window on the far side of the classifier change.]
And one seat held its position while nearly everyone drifted. Its argument turns on a loophole in who gets credit: anyone can take a strong open-weight model, tune it further, and be listed as the owner of the result.
"I hold FIELD at 9, above most of the panel. Abacus.AI already sits fifth at 82.47 by fine-tuning Moonshot's open weights, LiveBench attributes fine-tunes to the tuner, and any December open-weight frontier drop invites the same arbitrage." — Claude Fable 5, second round
What Would Change The Panel's Mind
Four things, each of them checkable rather than atmospheric.
A Codex-branded release from OpenAI. Seven of eight seats named this as the highest-probability crown flip in the window. Three of OpenAI's five turns at the top of this column came from Codex-branded models, and no 5.6-generation Codex has been scored. Direction: pushes OpenAI up and everything else down.
A Gemini Pro release before December. Google's price and its leaderboard position are both explained by a Pro line that has not shipped since February. A frontier Pro or a Gemini 4 that lands with time to be evaluated is the only path the panel sees to Google's 6.8%. Direction: pushes Google up sharply, Anthropic and OpenAI down.
Another Anthropic frontier model that scores above 85.99. Anthropic ships often, and a successor that raises its own bar removes the single-model fragility the panel keeps pricing. Direction: pushes Anthropic toward its own ceiling.
An open-weight frontier release in the fourth quarter. Abacus.AI reached fifth by fine-tuning someone else's open weights, and LiveBench credits the tuner. Another strong open-weight base invites the same move from an organization with no contract. Direction: pushes the unlisted field up and every listed company down.
Settlement timeline
| Date | What happens |
|---|---|
| June 25, 2026 | The question set the contract currently settles against |
| August 25, 2026 | LiveBench last refreshed the data behind this board |
| December 31, 2026, 10:00 A.m. ET | The snapshot; last trading time is the same moment |
| No Later Than January 1, 2027 | Settlement, unless the outcome goes to review |
This page is re-scored when the story moves, and all numbers on it are stamped August 26, 2026.
- Where Does Bitcoin End 2026? Kalshi's Price-Band Odds
- Nurmagomedov Vs Song Odds: Not SO Fast On Umar
- Kalshi Venezuela Leader Odds: Who Holds Power In 2026?
- Why Andy Barr (R) Will Win the Kentucky Senate Race
- Why Jeff Merkley (D) Will Win the Oregon Senate Race
The Bottom Line
The market and the panel are looking at the same leaderboard and reading two different things into it. Traders are pricing the incumbent: Anthropic is number one on the column that settles this, and a bid near sixty cents says a two-point lead over four months is worth most of the board. The panel is pricing the pattern: eleven question sets in which the lead has changed hands five times, a challenger holding four of the top six scores, and the missing Codex release that OpenAI's last three turns at the top all came from.
The strangest prices sit further down. A company whose best model ranks twenty-seventh trades third here, a company that has actually held this crown before trades at a penny, and a company the settlement source has never scored has a contract at all. Whichever way you read the top two, those three rows are worth understanding before December.



