There is a market on Kalshi that asks which company will have the number-one AI model before 2027. Twelve companies have a tradeable contract on it. The company sitting at the top of the leaderboard that decides the whole thing is not one of them.
That is the first strange fact about this board. The second is that the company traders like best, OpenAI at 33.5 cents, has a best model ranked 22nd on that same leaderboard. We handed the board to eight AI models, showed them the leaderboard and the binding contract terms, withheld every price, and asked each one to put a number on all twelve companies. They came back below the market on eleven of them.
The Quick Answer
Kalshi's 2026 top-AI board settles on the LMArena text leaderboard with the "Remove Style Control" toggle applied, and on that exact view the five models sharing the top rank all belong to Anthropic, which has no contract here. Our eight-model panel puts Meta highest among the twelve companies you can actually trade, and prices OpenAI at roughly a third of where the market has it.
The reason sits in one sentence of the contract terms, and it is the most important thing on this page. It comes next, with the full twelve-company board and the seats that changed their minds.
Free: The Weekly PM Market Brief — the 8-model panel's graded record, the week's biggest market-vs-model gaps, and what's spiking next. One email, Sundays.
New to Kalshi event markets? Our beginner's guide to how these contracts work covers the mechanics in about five minutes.
One thing to get straight before the numbers. Kalshi runs a second AI market asking who is on top on one specific day, December 31. We priced that board separately, and it is a different contract with a different answer: Claude trades at 69 cents there and ChatGPT at 13. Both numbers matter here, because the comparison between the two boards does most of the work later on.
The Rule Nobody Reads
Every contract here carries the same one-line summary: "If [company] has a #1 ranked AI model before Jan 1, 2027, then the market resolves to Yes." Simple enough. The binding terms, published by Kalshi as a contract-terms document, say three things that one line leaves out.
It is not the ranking you see first. The terms name "the Overall Arena Elo rankings (UB)" as the thing being measured. Elo is the leaderboard's rating number, built from millions of head-to-head votes between two anonymous models. UB is the upper-bound rank, and it counts only the models that are provably better: a model's UB rank is one plus the number of models whose worst plausible score still beats its best plausible score. Models too close to separate share a rank. A model can sit eighth in the visible list and third in the UB column, and the UB column is the one that counts. On top of that, Kalshi's own note on the market tells you to check the source with the "Remove Style Control" toggle applied. Style control is a statistical adjustment that strips out the effect of answer length and formatting, and turning it off changes who is at the top. Hold that thought, because it is where Meta's whole case lives.
It pays on any single day. The terms say the ranking is "checked at 10:00 AM ET daily." A company qualifies if it holds the top spot on any one of those checks, and its own contract stops trading that morning while the rest of the board carries on. That structure normally makes a contract worth more than a year-end snapshot, because you get every remaining morning instead of one.
And then the sentence that undoes all of it. From the payout criterion, word for word: "If they have an LLM that is tied with another LLM, then the Payout Criterion is not fulfilled." Reaching the top is not enough. A company's model has to be alone up there, statistically clear of every other model on the board.
Here is what that means today. On the settlement view, which is the text arena, overall, with style control removed, as of the leaderboard's August 21 vote cutoff across 7.9 million votes and 394 models: five models share the number-one UB rank, and all five are Anthropic's. Their raw scores differ slightly, from 1504.2 down to 1494.7; what they share is that none of them is provably better than the others, which is exactly the tie the contract describes. Nobody is alone at the top. Under a strict reading of that sentence, not one company on the planet qualifies right now, including the one winning.
The Board
Twelve companies, the market's price for each, and where each company's best model actually sits on the leaderboard that settles the contract.

| Company | Best model on the settlement view | UB rank | Market | Panel |
|---|---|---|---|---|
| OpenAI | gpt-5.5-high (row 22) | 14 | 33.5¢ | 12.3% |
| xAI | grok-4.5 (row 37) | 30 | 20.5¢ | 4.3% |
| Meta | muse-spark-1.2 xHigh (row 8) | 3 | 13.5¢ | 16.0% |
| Moonshot AI | kimi-k3-max (row 17) | 7 | 13¢ | 6.0% |
| Z.ai | glm-5.3-max (row 15) | 6 | 11¢ | 6.5% |
| Alibaba | qwen3.8-max (row 11) | 6 | 10¢ | 8.8% |
| Nvidia | nemotron-3-ultra-550b (row 51) | 33 | 8.5¢ | 1.1% |
| Deepseek | deepseek-v4-pro (row 40) | 32 | 8¢ | 3.8% |
| ByteDance | dola-seed-2.0-pro (row 43) | 34 | 7.5¢ | 1.5% |
| Baidu | ernie-5.1 (row 24) | 15 | 6.5¢ | 2.8% |
| Mistral | mistral-large-3 (row 84) | 71 | 4.5¢ | 0.9% |
| 01A1 | no model listed under this organization | n/a | 4¢ | 0.5% |
Market = the midpoint of the live bid and ask, meaning the highest price anyone is offering to buy at and the lowest anyone will sell at, fetched August 22, 2026 at 17:54 UTC. Six of the twelve quote a bid and an ask within five cents of each other; four are wider than seven cents. Panel = the median of eight independent model estimates after a revision round. Rows and UB ranks come from the settlement-view leaderboard, vote cutoff August 21, 2026. Kalshi's contract still says "xAI"; the leaderboard files Grok under SpaceXAI, the name the company took after SpaceX acquired xAI in an all-stock deal on February 2, 2026 and rebranded the combined AI unit in July. Model estimates, not predictions of fact and not financial advice.

Set the middle two columns side by side and the market's shape stops making sense. The three companies whose models sit closest to the top in the column the contract names, Meta at UB 3 and Alibaba and Z.ai at UB 6, trade at 13.5, 10 and 11 cents. The two companies trading highest, OpenAI and xAI, sit at UB 14 and UB 30. Traders have ranked this board close to backwards from the leaderboard it settles on.
More live boards from the same panel: Anthropic IPO odds, covering the leader you cannot trade here, with an IPO announcement before 2027 priced at 91¢ · who finishes 2026 on top, where Claude trades at 69¢ · what actually counts as an OpenAI IPO. Prices fetched August 22, 2026.
Every Seat's Number
Each seat saw the same data card: the full settlement-view leaderboard, both toggle views, the contract terms word for word, and every company's model count and vote totals. No prices at all. Then each seat read the other seven anonymously and could revise. These are the revised numbers.
| Company | Fable | Opus | Sonnet | GPT | Gemini | GLM | Kimi | DeepSeek | Median |
|---|---|---|---|---|---|---|---|---|---|
| Meta | 17 | 15 | 19 | 17 | 14.5 | 13 | 18 | 11 | 16.0 |
| OpenAI | 13 | 10 | 15 | 13 | 11.5 | 10 | 15 | 8 | 12.3 |
| Alibaba | 9 | 7 | 10 | 9 | 8.5 | 6 | 10 | 4.5 | 8.8 |
| Z.ai | 7 | 4.5 | 7 | 6.5 | 6.5 | 5 | 8 | 3.5 | 6.5 |
| Moonshot AI | 6 | 4.2 | 6 | 6 | 6 | 4 | 7 | 3 | 6.0 |
| xAI | 5 | 3.2 | 4 | 4.5 | 4.5 | 3 | 5 | 3 | 4.3 |
| Deepseek | 4 | 2.6 | 4 | 3.5 | 5 | 2.5 | 6 | 2.5 | 3.8 |
| Baidu | 3 | 1.7 | 3 | 2.5 | 3 | 2 | 3 | 1.5 | 2.8 |
| ByteDance | 2 | 1.2 | 1.5 | 1.5 | 2 | 1 | 3 | 1 | 1.5 |
| Nvidia | 1.5 | 1 | 1 | 1.2 | 1.5 | 0.5 | 2 | 1 | 1.1 |
| Mistral | 1 | 0.7 | 1 | 0.8 | 1 | 0.5 | 1 | 0.5 | 0.9 |
| 01A1 | 0.5 | 0.6 | 0.5 | 0.5 | 0.5 | 0.5 | 0.5 | 0.5 | 0.5 |
Every seat on this panel is graded against real market settlements — records to date: GPT 84% on 5,658 graded calls · Gemini 86% on 4,423 graded calls · GLM 81% on 2,437 graded calls · Kimi 82% on 2,414 graded calls · DeepSeek 80% on 2,462 graded calls. Recomputed daily; the full scoreboard is public.
These are model estimates, not predictions of fact and not financial advice.
Every number in this piece gets graded in public once the market settles, and the running tally lives on the full graded scoreboard.
There is a simple way to read the two price columns at once. Kalshi flags this board as not mutually exclusive, meaning one company qualifying does not stop another from qualifying later in the year. Each column can therefore be added up, and the total read as the number of companies expected to reach a sole number one before 2027. The market's twelve prices sum to 140.5 cents, so traders are collectively paying for about one and a half qualifiers. The panel's twelve numbers sum to 64.5, which is about two thirds of one. That single comparison is the disagreement, and everything below is the argument over it.
The Market Card
| Venue | Kalshi, a CFTC-regulated exchange for event contracts (18+; availability varies by state, as of August 2026) |
| The Board | KXTOPAI-27, twelve tradeable yes/no contracts, one per company, plus a thirteenth that has never traded |
| Settles On | The Overall Arena Elo ranking (UB) on the LMArena text leaderboard, style control removed, checked at 10:00 AM ET every day |
| The Catch | A company whose model is tied with another model does not qualify |
| Expires | January 1, 2027, or the first 10:00 AM ET after that company qualifies, whichever comes first. Each company's contract closes on its own qualifying day; the exchange flags the board as not mutually exclusive, so more than one can pay |
| Prices As Of | August 22, 2026, 17:54 UTC |
The books here are thin and uneven. Z.ai shows 2 cents bid against 20 asked, and Alibaba shows 5 against 15. On those two the round trip costs more than the position is plausibly worth, and the midpoints in the board above are doing a lot of work. Meta and Nvidia, by contrast, are quoted a single cent wide. The spread on screen, not the midpoint, is what any position actually costs.
That thirteenth contract is worth thirty seconds. Kalshi lists a separate market for "Zhipu AI" that has never traded a single contract and sits at zero. Zhipu AI rebranded itself as Z.ai internationally in July 2025, which means the board carries two contracts for one company: one at 11 cents with an open book, and one dead at zero.
Why The Panel Won't Pay 33.5 Cents For OpenAI
OpenAI is the most-traded name here and the panel's second choice, and those two facts sit a long way apart. Its best model on the settlement view is gpt-5.5-high at row 22, UB 14, with an Elo score 33.3 points below the leader across 59,545 votes. That vote count matters. A model with sixty thousand votes has a narrow confidence interval, which is another way of saying its rating is well established and has stopped drifting. Whatever OpenAI does on this board, it does with a new model.
The panel still gave it the second-highest number, and the seats agreed on why: of the twelve, OpenAI has the strongest history of shipping a release that jumps a long way at once, and it puts everything on this leaderboard. There are 44 OpenAI models on it, more than any other company on the entire 394-row leaderboard. The panel's 12.3% is a forecast about OpenAI's next model, not its current one.
Shipping is the easy part. The contract pays only for shipping something that stands alone above five Anthropic models, which is a much narrower event than a big release. The GPT seat, which had priced OpenAI at 19 in the first round and cut it to 13 after reading the others, put the whole board in one line.
"The board is shaped by strict sole-UB-1 settlement risk, with Meta closest today and OpenAI retaining the largest latent release-jump upside." — GPT
Then there is the number from the top of this page. On Kalshi's other top-AI market, the one asking who is on top on December 31, ChatGPT trades at 13 cents and Claude at 69. Same exchange, same underlying question, same year. OpenAI is more than two and a half times as expensive on the board where Anthropic has no contract as it is on the board where Anthropic does.
Meta Is The One Name The Panel Rates Above Its Price
Meta is the one company on this board the panel rates above its price, and the reason is a toggle.
On the settlement view, muse-spark-1.2 (xHigh) sits at row 8 with a UB rank of 3, an Elo score 16.8 points behind the leader, on just 3,257 votes. Flip style control back on, which is the leaderboard's default public view and not the one that settles these contracts, and the same model moves to row 4 with a UB rank of 1, tied at the top with three Anthropic models. Meta is already number one by the ranking Kalshi told you not to use.
A tie pays nothing, and the panel was careful about that. What the toggle shows is how narrow the gap is: one release, or one methodology decision, stands between Meta and the top. Meta also ships fast, with 21 models on the leaderboard and a muse-spark line that has gone from muse-spark to 1.1 to 1.2 in short order. A price of 13.5 cents implies roughly a 1-in-7 chance; the panel's 16% is closer to 1-in-6.
The sharpest bear case came from inside the panel, and it is the best reasoning the whole run produced. Before it, one more thing has to be said about the leaderboard itself.
The Leaderboard Itself Is Contested
Everything above assumes the settlement source measures what it appears to measure. That assumption has a documented, published record against it, and on a contract that settles against this exact leaderboard, the record is part of the price.
In April 2025 a team from Cohere Labs, AI2, Princeton, Stanford, Waterloo and the University of Washington published The Leaderboard Illusion, which argues the arena's evaluation framework is distorted by undisclosed private testing. Its central finding is that large providers can test many variants privately and publish only the best result. The paper counts "27 private LLM variants tested by Meta in the lead-up to the Llama-4 release," and estimates that Google and OpenAI received roughly 19.2% and 20.4% of all arena data respectively, while 83 open-weight models shared about 29.7% between them. It also estimates that even limited additional data can produce relative performance gains of up to 112% on the arena distribution.
A second paper, Improving Your Model Ranking on Chatbot Arena by Vote Rigging, published in January 2025, goes at the votes themselves. Running experiments on about 1.7 million historical arena votes, its authors conclude that "omnipresent rigging strategies can improve model rankings by rigging only hundreds of new votes," because under an Elo system any new vote can move a target model's rank even when that model is not in the battle.
Neither paper is an accusation that this specific board is being manipulated, and neither changes what the contract says. Both matter to the price for the same reason: the thing being measured is a crowd vote with a documented history of being steerable, and the number of votes behind a model varies enormously across this board.
| Company's Best Model | Votes behind the rating | UB rank |
|---|---|---|
| Meta, Muse-Spark-1.2 xHigh | 3,257 | 3 |
| Z.ai, Glm-5.3-Max | 3,751 | 6 |
| Alibaba, Qwen3.8-Max | 9,955 | 6 |
| Moonshot, Kimi-K3-Max | 15,054 | 7 |
| xAI, Grok-4.5 | 22,030 | 30 |
| OpenAI, Gpt-5.5-High | 59,545 | 14 |
Vote counts from the settlement-view leaderboard, August 21, 2026 cutoff.
A rating built on three thousand votes moves for reasons a rating built on sixty thousand does not, and the contract pays on a single morning's reading. The two best-placed challengers in the UB column, Meta and Z.ai, are also the two whose ratings rest on the thinnest evidence.
The Meta finding cuts both ways, which is why the panel did not treat it as a reason to sell. A company willing to test 27 variants to find its best one is a company optimizing hard for exactly this leaderboard, and Meta's price is the one the panel already likes.
Where The Panel Changed Its Mind
After round one each seat read the other seven anonymously and could revise. Three moved hard. One barely moved at all, and that is the one worth reading first.
Opus held, and explained why the others had the mechanism backwards. Several seats had built a bull case on the debut window: the idea that a brand-new model with very few votes carries a wide confidence interval, which could briefly push its upper bound above everyone else's and win the contract for a morning before the votes narrow it. Opus took that apart, and moved its own numbers by a point or two in either direction.
"Sole rank(UB) 1 requires your lower bound above everyone's upper bound, so wide-CI debuts can't win and multi-variant labs tie themselves." — Opus
In plain terms: having few votes makes a model look potentially great and potentially terrible at the same time. That gets it into the tie at the top, which pays nothing, and it actively prevents the model from clearing that tie. The winning shape is the opposite, a big rating jump plus enough votes to prove it. The second half of that line is the sharper half. A lab that submits many models is partly working against itself, because its own entries crowd the top band. Anthropic's top two models sit 0.2 Elo points apart, which is exactly why nobody qualifies today.
GPT pulled its whole board down. Meta went from 31 to 17, the largest single revision of the run, and Alibaba from 17 to 9, once it recalculated what the UB column and the tie clause do together.
Gemini cut OpenAI almost in half, from 25 to 11.5. It had been the most bullish seat on OpenAI in round one. Its revised reasoning named the same bar:
"The strict tie clause requires challengers to statistically clear a dense Anthropic cluster, leaving low-vote debut spikes (Meta) or massive step-change releases (OpenAI) as the only paths to victory." — Gemini
[Editor's note: the first of those two paths is the one Opus argued against in the passage above. A low-vote debut widens a model's interval in both directions, which is what stops it from clearing the tie. We publish the seat's reasoning as it was given, and the disagreement between the two seats is itself part of the record.]
Kimi went the other way, raising OpenAI from 7 to 15 and Meta from 12 to 18, on the argument that a handful of step-change releases happen every quarter across the industry and there are a lot of mornings left to catch one.
The spread across the eight seats on OpenAI fell from 18 points to 7 over the revision round, and on Meta from 19 to 8. That is a larger convergence than this panel usually produces.
Dark Horses The Panel Won't Dismiss
The top two names take most of the attention. Three companies underneath them have better current standing than their prices suggest.
Alibaba, 10¢, panel 8.8%. qwen3.8-max sits at UB 6, eight UB places closer to the top than OpenAI's best model, and Alibaba has 42 models on the leaderboard, second only to OpenAI's 44. The GPT seat marked it down on the judgment that a single Qwen release is less likely to separate cleanly, landing at 9 against a 10-cent market, the closest any seat came to meeting a price on this board apart from Meta.
Z.ai, 11¢, panel 6.5%. glm-5.3-max is at UB 6 on only 3,751 votes. The GLM seat, pricing its own maker's contract, was one of the more cautious voices on it rather than one of the most bullish. Note the book, though. Two bid against twenty asked is barely a quote.
Moonshot AI, 13¢, panel 6.0%. kimi-k3-max sits at UB 7 with 15,054 votes, a firmer position than Z.ai's on four times the evidence. Moonshot has only 7 models on the leaderboard, which is fewer shots on goal than the labs around it.
Three more the panel effectively wrote off. Nvidia trades at 8.5 cents with a best model at row 51 and a UB rank of 33, and every seat priced it at 2% or below. Mistral trades at 4.5 cents from row 84. ByteDance trades at 7.5 cents with exactly one model on the entire 394-row leaderboard, at UB 34. The DeepSeek seat, pricing its own maker's contract at 2.5%, was the most bearish voice about the board as a whole:
"The tie clause requires sole UB=1 separation above five clustered Anthropic models, making this a clean-step-change contract where only Meta's proximity and OpenAI's release history justify >5% probabilities." — DeepSeek
By that standard, ten of the twelve contracts on this board are trading above the highest number any seat gave them. Only Meta and Alibaba have a seat willing to meet the market.
The Two Companies You Cannot Trade Here
Anthropic holds all five of the number-one UB slots on the settlement view, led by claude-opus-5-high at 1504.2 Elo. Google's gemini-3.7-flash-high is the highest-placed model from any other company, at row 6 and UB 3. Neither has a contract on this board.
Kalshi does not publish its listing decisions, so why they are absent is not something we can tell you. What we can tell you is that the exchange prices the same fact elsewhere. On Kalshi's weekly top-model market (KXTOPMODEL-26AUG24), claude-opus-5-high is quoted at 91 bid against 94 asked to be the top model this week. Traders on that board are near-certain about who is winning; traders on this one are paying 33.5 cents for the company that is 14th.
The effect is knowable even if the reason is not. Every contract here prices one of twelve challengers dislodging a company that is not on the menu. Eleven of these contracts opened on January 1 and ByteDance's opened in February, and not one has had a qualifying day since.
What Would Change The Panel's Mind
Four things, each checkable, each with the direction it pushes.
- Meta ships the next muse-spark and it lands materially above 1504 Elo with enough votes to tighten its interval. Pushes Meta sharply up and everything else down, because a qualifying day closes Meta's contract and proves the tie clause is beatable. This was the panel's clearest single trigger.
- OpenAI ships a new frontier model and it lands inside the Anthropic pack instead of above it. Pushes OpenAI down hard, and would be the strongest evidence yet that 33.5 cents was paying for a release rather than for a win.
- Anthropic or Google puts a model alone at the top. Pushes every tradeable number down. The wall gets higher and the untradeable labs absorb the outcome, which is the scenario the DeepSeek seat was pricing.
- The number-one UB group shrinks from five models to one through vote drift alone. Whichever way that goes, it produces the first public evidence of how the tie clause behaves in practice, currently the largest unpriced risk on the board. Every seat named it as the thing it was least sure about.
Settlement Timeline
| Checked | Every day at 10:00 AM ET, on the style-control-removed LMArena text leaderboard |
| Next Catalyst | Any frontier release from the twelve listed companies; new models were being added to the leaderboard near-daily through August 2026 |
| Second Catalyst | Movement in the five-model tie at UB 1, which can happen on vote drift alone, with no release at all |
| Third Catalyst | The December 31, 2026 reading of Kalshi's separate year-end top-AI board, which resolves the same question on a single date |
| Last Trading | The first 10:00 AM ET after that company qualifies, or 11:59 PM ET on December 31, 2026 |
| Expires | January 1, 2027 |
This page is re-scored when the story moves. Prices and leaderboard positions are as of August 22, 2026.
The panel's calls are on the record. Our 8-model AI panel puts a graded, public verdict on the December 31 version of this same question, so if you want the year-end view rather than the any-morning view, the live verdict on who finishes 2026 on top has it.
- BMW Championship Top 5 Odds 2026: Kalshi Odds
- Golf Top 20 Parlay Odds: Scheffler Plus Two: Kalshi Odds
- FedEx St Jude Top 5 Odds On Kalshi: Same Score, 3x The Price
- FedEx St. Jude Top 20 Odds: The Kalshi Miss We Called First
The Bottom Line
The gap the panel found here is a disagreement about what the contract asks, more than a disagreement about which lab is best. The summary line on the market page describes a contract on who builds the strongest AI, with a favorite you can name and a lot of mornings for it to happen. The terms describe a contract on one model standing statistically alone above every other model on a 394-row leaderboard, on a specific morning, measured by a column most people never look at, on a view of the data you have to switch on deliberately.
Those are different propositions, and on our reading of the tie clause the second is much harder. Eight models read the same terms independently and landed in the same place: on eleven of these twelve contracts the price is paying for the first description, while the second one is what settles.
The one name the panel's numbers leave standing is Meta at 13.5 cents, sitting three UB places from the top and already tied for first under the other toggle. That is the whole edge the panel found, and it is a small one on a thin book. The arithmetic changes with the spread on screen, which on this board is often wide.



