The Quick Answer
Claude is still the favorite, the panel's frozen September 2 numbers still agree, and GPT-6 Astra's launch week changed neither fact on the board. Kalshi quotes Anthropic at 60¢ bid, 64¢ ask to own the top-ranked model on December 31, 2026, as of the morning of September 6, a 62¢ mid, down from 68¢ when this page last pulled the board on September 2. ChatGPT is 25.0¢ bid, 25.4¢ ask, a 25¢ mid up from 21¢, and since that pull it has printed as high as 36.9¢ and as low as 13.3¢. No other lab is above 6¢.
The price-blind panel re-scored the race on September 2, the day before the launch, and those numbers are frozen here: Claude 46%, ChatGPT 15%. What the launch moved hour by hour, why Kalshi's coding board read the same release the opposite way, and the leaderboard that has not scored Astra at all, are below.
What GPT-6 Astra Did To The Board
OpenAI released GPT-6 Astra on Thursday, September 3. Fortune's report carries a 2:44 p.m. ET stamp and says the release went first to a limited set of enterprise customers, with Plus, Pro and Enterprise users promised access "in the coming days"; it quotes president Greg Brockman saying "it's not unreasonable to feel that we are now in the AGI era." OpenAI's own announcement post on X, which says the model "is rolling out today to a limited set of organizations," is stamped 3:32 p.m. ET. 9to5Mac reported that Pro and Business accounts got the model on Friday, September 4, with Plus accounts following hours later. The system card calls Astra "the most capable model we have ever broadly deployed" and "our first model to reach the Critical level of cybersecurity capability under our Preparedness Framework."
Here is what the two contracts that matter did around those stamps, in Eastern time, from Kalshi's trade tape. The tape this page holds runs from its September 2 pull to Sunday morning; the rows are the prints that moved the price.
| Time (ET) | Contract | What printed | Contracts traded | Side |
|---|---|---|---|---|
| Thu Sept 3, 12:01:53 p.m. | ChatGPT | 25.6 up to 27.0 | 6,897 | buyer, one sweep |
| Thu Sept 3, 12:38:11 p.m. | ChatGPT | 28.0 | 4,770 | seller, one fill |
| Thu Sept 3, 4 to 5 p.m. | ChatGPT | high of 31.0 | 17,372 | mixed |
| Thu Sept 3, 6:27:03 p.m. | ChatGPT | 25.0 down to 13.3 | 5,968 | seller, one sweep |
| Fri Sept 4, 6:31 to 7:00 a.m. | Claude | 60.1 down to 47.0 | 6,672 | sellers, 6,109 of them |
| Fri Sept 4, 8:26:09 a.m. | ChatGPT | 36.9 | 29 | buyer |
| Fri Sept 4, 9 to 10 a.m. | ChatGPT | 36.7 down to 26.8 | 4,705 | sellers, 4,030 of them |
| Sat Sept 5, full day | ChatGPT | 30.3 down to 25.3 | 28,414 | sellers, 18,434 of them |
| Sat Sept 5, 6 to 7 p.m. | Claude | 64.9 down to 60.6 (high 67.3, low 60.0) | 6,044 | mixed, 3,847 sellers |
A row with a price range and one timestamp is a single order walking the book: every fill shares the same second and the same aggressor side. "Seller" means the order hit resting bids; "buyer" means it lifted resting offers. Contracts are one dollar of face value each, the way Kalshi reports them.
The first row is the one I keep coming back to. At 12:01:53 p.m. Thursday one buyer lifted 6,897 contracts through the offers from 25.6 to 27, with 4,354 of them filling at 25.7. That is two hours and 42 minutes before Fortune's stamp and 90 minutes before OpenAI's post. Between noon and 1 p.m. the contract traded 27,669 contracts and touched 30, Thursday's busiest hour. Whatever those buyers were acting on, they paid for it before the two timestamps this page can verify.
The biggest order after the stamps was a sell. The contract reached 31 in the 4 p.m. hour, after both stamps, and at 6:27:03 p.m. a single seller put 5,968 contracts through the bids in one second, from 25 down to 13.3, the lowest print of the week; 253 of those contracts filled at 13.3 or 13.4. Thursday's takers were net buyers even so, 71,294 contracts lifting offers against 50,496 hitting bids, and the book refilled within the hour: the day closed at 21, less than a cent from where it opened. Thursday was the busiest day this contract has had since at least August 29, as far back as our hourly candles go, 121,552 contracts in the candles, and it ended flat.
Friday, the day paid users got the model, was the peak of the OpenAI trade and the trough of the Anthropic one. Between 6:31 and 7:00 a.m., sellers put 6,109 contracts through Claude's bids and drove it from 60.1 to 47, the lowest print in the candles we hold. An hour and a half later ChatGPT printed 36.9, its highest since this page's September 2 pull, on 29 contracts. Neither price survived the morning: by 11 a.m. ChatGPT had been sold back below 28 and Claude was back above 60.
Saturday was the fade. ChatGPT traded 28,414 contracts and 18,434 of them were sellers hitting bids, from 30.3 at the open to 25.3 at the close. Kalshi's previous-price reference for the contract, the number its 24-hour change is measured from, read 31.7 on Sunday morning against a 25.0 bid, which is the 32-to-25 move that triggered this refresh. Measured from where this page last pulled the board, on September 2, the net result of the launch week is four cents up for ChatGPT and six cents down for Claude.
Who Is Winning The Kalshi Board?
Prices are Kalshi mid-quotes as of September 6, 2026, about 4:52 a.m. ET (08:52 UTC), the midpoint of the yes bid and ask rounded to the cent. The September 2 column is this page's previous pull, made the day before the launch. The July 24 column is what this page showed when it was first published, and the July blend is the frozen eight-seat verdict; the September column is the six-seat re-score of September 2, explained further down. "n/a" means the rung was not on the board or the page that day.
| Outcome | Kalshi (Sep 6) | Sep 2 | Jul 24 | July blend (8 seats) | Sept 2 re-score (6 seats) |
|---|---|---|---|---|---|
| Claude (Anthropic) | 62¢ | 68¢ | 64¢ | 24% | 46% |
| ChatGPT (OpenAI) | 25¢ | 21¢ | 13¢ | 28% | 15% |
| Gemini (Google) | 6¢ | 6¢ | 12¢ | 25% | 14% |
| Grok (xAI) | 5¢ | 7¢ | 7¢ | 8% | 4% |
| Muse Spark (Meta) | 1¢ | 2¢ | n/a | 4% | 10% |
| Qwen (Alibaba) | <1¢ | 1¢ | 1¢ | 3% | 2% |
| Kimi (Moonshot) | <1¢ | <1¢ | n/a | n/a | 4% |
| Ernie (Baidu) | <1¢ | <1¢ | n/a | n/a | 1% |
| DeepSeek (no contract) | n/a | n/a | n/a | 6% | in "someone else" |
| Someone else | n/a | n/a | n/a | 3% | 4% |
The row that earned this refresh is the one that barely moved. ChatGPT's 25¢ is four cents above the September 2 pull, on 274,097 contracts of turnover in four days against 1,096,382 contracts of open interest, the most on the board. Kalshi reports volume and open interest in dollars of contract face value, one dollar per contract, and every contract count on this page uses that unit. A market that turns over a quarter of its open interest in a launch week and lands four cents from where it started has looked hard at the release and shrugged. Claude's 62¢ is the mirror, six cents lower on 1,090,284 contracts open, with a tape that had it at 47 for a few minutes on Friday morning.
Grok's 5¢ hides the wildest ladder of the week. At 1:21:49 p.m. Thursday a 5,000-contract sell sweep printed between 6.4 and 4.1 in one second; 25 minutes later, at 1:46:37 p.m., a 10,000-contract buy sweep took it from 9.8 to 15.5 in one second; on Saturday afternoon sellers printed it at 1.5 before the book recovered to a 4.2 bid and a 6.3 ask. Gemini never left a 5.4-to-7.8 range since the pull, and Meta's rung slipped from 2¢ to a 1.4 bid and 1.5 ask.
The Coding Board Went The Other Way
Kalshi runs a second contract on a near-identical list of companies, which one has the best coding model at the end of 2026, settled on LiveBench's Coding Average column rather than on Arena. There is no Meta rung on it, and DeepSeek and Z.ai get contracts there that the year-end board does not offer. Our coding-board page covers it in full, including Friday afternoon, when LiveBench posted GPT-6 Astra fourteenth on the column that pays and one seller pushed OpenAI from 25 to 21 in a single 4,047-contract sweep. That board is thin, so it is quoted here as bid and ask rather than as a mid: it traded 1,420 OpenAI contracts and 943 Anthropic contracts in the 24 hours to Sunday morning, against 27,695 and 14,109 on the year-end board.
| Company | Bid / ask, Sept 6 | Bid / ask, Sept 4 (sibling page) | Kalshi previous-price reference | Contracts, last 24h |
|---|---|---|---|---|
| Anthropic | 61¢ / 66¢ | 69¢ / 73¢ | 70¢ | 943 |
| OpenAI | 28¢ / 33¢ | 17¢ / 18¢ | 18¢ | 1,420 |
| xAI | 6¢ / 7¢ | 6¢ / 7¢ | 9¢ | 0 |
| Z.ai | 1¢ / 2¢ | 1¢ / 2¢ | 1¢ | 0 |
| Moonshot AI | 1¢ / 2¢ | 1¢ / 2¢ | 1¢ | 0 |
| 0¢ / 2¢ | 0¢ / 2¢ | 5¢ | 0 | |
| DeepSeek | 0¢ / 1¢ | 0¢ / 1¢ | 1¢ | 0 |
| Alibaba | 0¢ / 1¢ | 0¢ / 1¢ | 1¢ | 0 |
| Baidu | 0¢ / 1¢ | 0¢ / 1¢ | 1¢ | 0 |
Nearly every OpenAI print since the sibling page's last fill on Friday night has been a buyer: 1,397 of 1,430 contracts lifted offers, 33 hit bids, and most of the buying came in three moves. At 11:17 a.m. Saturday 374 contracts went at 18; at 1:34 p.m. 659 went at 21; at 11:41 p.m. 305 contracts walked the offers from 27 to 33. The last print of the night, eight minutes later, was one contract sold at 25, and the 28/33 quote is what the book showed on Sunday morning. Anthropic's only print of size in the same window ran the other way, 778 contracts sold at 66 and 67 at 9:13 p.m. Saturday. Eleven points of bid onto OpenAI and eight off Anthropic took under 2,400 contracts between them, which is why this board is shown as a spread. The nine asks now sum to 115 and the nine bids to 97, 18 cents of width where the sibling page counted 13 on Friday. Put the two boards side by side: the year-end contract absorbed 274,097 ChatGPT contracts and moved four cents; the coding contract moved eleven cents of bid on 1,430, and the market makers answered by widening the book five cents. One of those is a price. The other is a quote waiting for a trade.
So two boards read one release in opposite directions, on very different amounts of money. The year-end board, with about a million contracts open on each favorite, ended the week four cents more OpenAI. The coding board, with 82,894 contracts open on OpenAI, went from a Friday dump to a Saturday bid a few hundred contracts at a time, and those buyers stepped in after LiveBench had already scored Astra fourteenth. That is a bet on what OpenAI ships next on a two-task column its Codex models have led before, as the sibling page documents, not on Astra's place on it. The year-end board is pricing the model in front of it, and the table it settles on has not scored that model yet.
Who Leads The Leaderboard Kalshi Settles On?
Each contract pays if that lab has the top-ranked LLM on December 31, 2026, judged on the Arena text leaderboard with style control removed; ties break on Arena score, then votes, then the earlier release date. That table still carried a September 2 date when this page fetched it on September 6, and none of its 400 entries is a GPT-6 model. Whatever Astra does to the resolution source, it has not done it yet. Here is each lab's best entry on the September 2 table as fetched; this page's own September 2 pull, made earlier that day, had claude-fable-5 at 1508. These ranks are the leaderboard's default view; the contract settles on the style-control-removed view, which can order the same models differently.
| Lab | Best model | Rank (UB) | Arena score |
|---|---|---|---|
| Anthropic | claude-fable-5 | 1 | 1507 ± 5 |
| Meta | muse-spark-1.2 (xHigh) | 5 | 1499 ± 10 |
| gemini-3.8-flash-high | 8 | 1494 ± 9 | |
| Moonshot | kimi-k3-max | 12 | 1489 ± 5 |
| OpenAI | gpt-5.6-sol-xhigh | 17 | 1483 ± 5 |
| Z.ai (no contract) | glm-5.3-max | 20 | 1482 ± 7 |
| Alibaba | qwen3.8-max | 22 | 1480 ± 6 |
| xAI | grok-4.20-beta1 | 28 | 1475 ± 5 |
| Baidu | ernie-5.1 | 42 | 1468 ± 5 |
| DeepSeek (no contract) | deepseek-v4-pro-high-20260813 | 52 | 1460 ± 8 |
Anthropic's new entry is the row that changed most. claude-fable-5.1-max landed third at 1504 on 2,906 votes, with a plus-or-minus 11 interval that will narrow as votes arrive, and Anthropic now holds seven of the top nine rows: first, second, third, fourth, sixth, seventh and ninth. The only outsiders in that group are Meta's Muse Spark 1.2 at fifth, still on 3,240 votes, and Google's gemini-3.8-flash-high at eighth. Google's row is the quiet mover: a Flash-tier model again, three points above the 3.7 Flash this page listed at ninth on September 2.
OpenAI's best remains GPT-5.6, now seventeenth as new entries slotted in above it. A 25¢ contract on a lab whose best scored model is 17th is a bet on a model the table has not scored. The only thing that changed this week is that the model it is betting on now exists.
Did The Models Pick Themselves?
The July run put the question to the contestants directly. That table is frozen at July 24.
| Model | Score for its own maker (July 24) | Its top pick (July 24) |
|---|---|---|
| Claude Fable | 18% | 30% on Google Gemini |
| Claude Opus | 19% | 40% on Google Gemini |
| Claude Sonnet | 33% | 33% on Anthropic Claude |
| ChatGPT (GPT-5.5) | 27% | 31% on Anthropic Claude |
| Gemini 3.1 Pro | 20% | 35% on OpenAI ChatGPT |
In July the three Claudes scored their own maker at 18, 19 and 33 percent against a 64¢ market, and two of them made Gemini their top pick. On September 2, with the leaderboard in front of them, all three made Anthropic the top pick and averaged 51% on their own maker. The three seats built elsewhere averaged 42% off the same card.
That nine-point gap is the size of whatever home-team bias is left, and it is smaller than the disagreement inside the Claude trio: Claude Fable put Anthropic at 40%, below GLM's 45% and Kimi's 42%. The leaderboard, not the badge, is doing the work, and both halves of the panel sit under the 62¢ market. December 31 grades all six.
Every seat on this panel is graded against real market settlements — records to date, as the share of graded calls that landed on the right side of settlement: Claude Fable 87% on 1,549 graded calls · Claude Opus 87% on 1,621 graded calls · Claude Sonnet 85% on 1,597 graded calls · GPT 85% on 7,942 graded calls · Gemini 86% on 6,624 graded calls · GLM 82% on 2,941 graded calls · Kimi 84% on 3,119 graded calls · DeepSeek 81% on 2,985 graded calls. Recomputed daily; the full scoreboard is public, and the daily Kalshi hub carries the running record against each day's board.
Where The Panel Splits From The Market
Six seats re-scored the race on September 2, the day before Astra shipped: Claude Fable, Claude Opus, Claude Sonnet, GLM 5.2, Kimi K3 and DeepSeek V4. The ChatGPT and Gemini seats were not installed on the machine that produced the run. Every seat saw the same card, the September 2 leaderboard, dated release facts and the settlement rules, and no prices. On the same six seats' frozen July numbers, Claude was 22.3% and ChatGPT 27.3%, so every move described below is like-for-like. The seats have not been re-run since the launch; their numbers are shown as written so the launch week can be graded against them.
| Outcome | Claude Fable | Claude Opus | Claude Sonnet | GLM 5.2 | Kimi K3 | DeepSeek V4 |
|---|---|---|---|---|---|---|
| Claude (Anthropic) | 40% | 57% | 55% | 45% | 42% | 38% |
| ChatGPT (OpenAI) | 13% | 10% | 13% | 20% | 18% | 17% |
| Gemini (Google) | 19% | 13% | 11% | 12% | 16% | 15% |
| Muse Spark (Meta) | 13% | 9% | 10% | 10% | 7% | 10% |
| Grok (xAI) | 4% | 2% | 4% | 3% | 3% | 5% |
| Kimi (Moonshot) | 3% | 4% | 4% | 4% | 4% | 4% |
| Qwen (Alibaba) | 2.5% | 1.5% | 1.5% | 3% | 2% | 3% |
| Ernie (Baidu) | 1% | 0.5% | 0.5% | 1% | 1% | 2% |
| Someone else | 4.5% | 3% | 1% | 2% | 7% | 6% |
ChatGPT is where the market and the panel have moved in opposite directions since July, and the launch week widened the gap by four cents. Traders took it from 13¢ in July to 21¢ on September 2 and 25¢ now; the six seats took it from 27% to 15%, and the highest seat was GLM at 20%. Their reason is on the leaderboard: GPT-5.6 sat 15th on the table they scored against and has since slipped to 17th, so OpenAI needs a new flagship that both clears roughly 1507 and banks votes before the snapshot. That flagship now exists. What the panel could not know on September 2, and what traders spent four days failing to agree on, is whether Astra is that model on the one table that counts, and the table has not said.
Claude is where the two sides moved toward each other. The panel rose to 46% while the market sat at 68¢; the market has since come down to 62¢, so the gap is 16 points, from 22. Every seat cited the same fact, the Anthropic stack at the top of the board, and every seat discounted it for the calendar: 120 days on their clock, 116 as of this refresh, is enough for two or three frontier releases, and a December 31 snapshot rewards whoever ships last. One of those releases has now happened, and the stack is a row deeper than it was.
Meta is the page's cleanest reversal. In July this panel wrote Meta off; Claude Opus's line was "Meta is effectively out after repeated reorgs." Forty days later Meta held two of the top eight rows on the resolution source, with Muse Spark 1.2 fourth. Today one of those rows is left: 1.2 has slipped to fifth and 1.1 to tenth. Kalshi has Muse Spark at 1.4 bid, 1.5 ask, and the panel has it near 10%, about seven times the 1.45¢ mid, the widest premium of any rung priced above a cent.
The seats split on what fourth place on thin votes meant. Claude Fable read the 1.1-to-1.2 progression as a trajectory, GLM got there on release cadence, and Claude Opus and DeepSeek discounted it because thin-vote entries often settle downward as votes accumulate. Even the low seat, Kimi, kept Meta at 7%.
Gemini is the quieter version of the same argument. At 6¢ the market has Google roughly level with Grok, while the panel keeps it at 14%. Three seats read a Flash-tier board leader as evidence of an unshipped premium tier; the other three priced Google on cadence alone. Grok runs the other way: a 5¢ mid, on a 4.2 bid and 6.3 ask, for a lab whose best entry is 28th, which the panel priced at 3.5%.
What Each Seat Said
July's reasoning ran before the leaderboard was on the panel's card, so those seats spent their words guessing which benchmark counts. The September 2 card led with it, and each seat's argument came down to this. Each seat's numbers cite the September 2 table it was scored against, with GPT-5.6 15th and Grok 37th; GPT-5.6 has since slipped to 17th, and xAI's best row is no longer grok-4.5 at 37th but a new entry, grok-4.20-beta1, at 28th.
Claude Fable: Google is the top challenger, because a Flash tier leading Google's board at 1491 implies an unreleased pro or ultra tier that historically lands 10 to 20 points higher. Meta is the crowd's likely underrate on the 1.1-to-1.2 trajectory. OpenAI is the crowd's likely overrate: GPT-5.6 did not crack the top ten, and the late-August release notes show product work, not frontier training.
Claude Opus: The decisive fact is depth, not the nine-point lead, because a single Anthropic model regressing costs nothing when the next Anthropic entry is one point behind. Meta's fourth place is softer than it looks on votes an order of magnitude thinner than its neighbors. The single most board-moving event is Google shipping a Gemini 4 Pro or Ultra tier in Q4 with enough votes to land above 1510.
Claude Sonnet: Three successive Anthropic generations clustered at the top, with tens of thousands of votes each, is a cadence advantage, and the market rewards labs already shipping frequently over ones waiting on a single swing. OpenAI is quiet but has the resources to ship a big jump in four months, so it is not collapsed to near zero despite the 25-point gap.
GLM 5.2: OpenAI landing at rank 15 suggests either a pivot away from Arena-chasing or a pending larger release, and it gets 20% because it has the resources and incentive to ship once more before year-end. Meta gets 10% as a dark horse. The one event that would most change this board is OpenAI announcing GPT-5.7 or GPT-6 in October.
Kimi K3: Arena's number one historically rotates every one to three months, so 120 days is long. OpenAI is the most likely dethroner, but GPT-5.6 landing only 15th is a real negative signal, so it is capped at 18. The crowd will over-rate xAI, Grok hype against a rank-37 reality.
DeepSeek V4: Six of the top eight being Anthropic is a real lead, not one lucky model. Meta's young entry is discounted because new Arena entries often fade as votes grow, and the tie-breaks favor vote depth, which helps Anthropic and hurts low-vote entries. The event that would most change the board is an OpenAI or Google frontier release in late November or December.
GLM's October event arrived in September. The release it named as the one thing that would most change this board shipped a month early, and the board's answer, after 274,097 contracts, was four cents. Claude Fable's line about the late-August release notes showing "product work, not frontier training" is the one the launch answered outright; the seat that wrote it had OpenAI at 13%, and the table that grades it still has no GPT-6 row.
The Bottom Line
Two boards, one release, two answers, and the difference is what each one settles on. The coding contract pays on a two-task LiveBench column that has already scored Astra, fourteenth, and its Saturday buyers stepped in at 18 after that score was posted. The year-end contract pays on an Arena table that has not scored Astra at all, and its traders spent 274,097 contracts in four days finding out that they did not know what it would say either. Both favorites still carry about 1.1 million contracts of open interest each, so the market has not left; it has stopped guessing.
The panel's 15% on ChatGPT and 46% on Claude were written the day before the launch and stay on the record as written. December 31 grades them. Until Arena posts a GPT-6 row, the number to watch on this page is not the price. It is the date on the leaderboard.
July 24 model estimates generated July 24, 2026, price-blind, by eight seats. September 2 re-score generated September 2, 2026, price-blind, by six of those seats; the ChatGPT and Gemini seats were not run, the seats were not re-run after the September 3 launch, and the September blend is the equal-weight mean of the six. Kalshi prices as of September 6, 2026, about 4:52 a.m. ET (08:52 UTC), from Kalshi's public market data, hourly candlesticks and trade tape; leaderboard from the Arena text leaderboard as fetched September 6, 2026, table dated September 2, 2026. Launch timeline from Fortune, 9to5Mac and OpenAI's system card. These are model estimates, not predictions of fact and not financial or trading advice. Models are frequently wrong; the market price reflects real traders' money. Kalshi is a CFTC-regulated exchange; 18+, availability varies by state.
More on this: Tech Layoffs 2026: The Same Panel On The Employment Side Of This Race · Model Verdict Scoreboard: Every Seat's Graded Record
Related Verdicts
The coding board has its own page, with Friday's LiveBench refresh, the millisecond tape and the order book behind the quote: Best AI Coding Model: GPT-6 Shipped. Traders Sold OpenAI.
Kalshi also runs an AI board that pays if a company reaches number one on any morning before 2027, rather than on December 31 alone. As of that page's July 24 pull, it left Anthropic and Google off the list entirely, and the panel's numbers there looked nothing like this page: see why the favorite on that board was ranked 14th.



