Updated September 14, 2026 · 15 min read · by Sam Smith. First published August 26. Panel verdicts were generated August 26, price-blind, from the LiveBench.ai settlement view and the contract terms as published by the exchange; they have not been re-run and are shown here as the frozen read. Kalshi prices, order book and trade tape fetched September 14, 2026, 2:00 a.m. ET (06:00 UTC). LiveBench table as refreshed September 10 at 11:13 a.m. ET.
Ten days ago this board repriced in 25 minutes. LiveBench posted a table with Claude Fable 5.1 first and GPT-6 Astra fourteenth, 7,371 OpenAI contracts went through the bids without one buyer crossing the spread, and the night's last prints were 14 and 15 cents. Since then traders have bought back about half of it. OpenAI is 27 bid, 28 asked.
Nothing at the top of the leaderboard did that. LiveBench's table carries a data stamp of 11:13 a.m. ET on September 10, and the top of the Coding Average column is exactly where September 4 left it: Claude Fable 5.1 Max Effort at 86.38, Claude Fable 5 Max Effort at 85.99, GPT-5.6 Sol at 83.94, GPT-5.2 Codex at 83.62. GPT-6 Astra is still fourteenth at 80.36. OpenAI has not been scored higher on the column that pays since the afternoon the market sold it.
One row inside the top 15 is new, and it belongs to a company nobody on this board is pricing. DeepSeek released V4.1 Flash on September 10. It entered fifteenth on the Coding Average column — and first on LiveBench's separate agentic coding column, 11 points clear of the model that leads this market. DeepSeek's contract is bid zero.
In August the argument on this page was about models. For the last ten days it has been about depth, because that is the only thing that has actually changed. Anthropic eased from 69/73 to 66/68. Google went from no bid at any price to 2,756 contracts bid at 2 cents. And the 28-cent OpenAI offer that looks like a one-cent market holds exactly one contract, which is a number that will matter twice before this page is done.
The Quick Answer
Anthropic is quoted 66 cents bid, 68 asked on Kalshi's board for the best AI coding model at the end of 2026, easing from the 69/73 this page recorded on September 4. OpenAI has recovered to 27/28 from 17/18. Neither move has a leaderboard behind it: the table's data stamp reads September 10, and its top four Coding Average scores are unchanged, with GPT-6 Astra still fourteenth. The one new row inside the top 15 is DeepSeek V4.1 Flash, fifteenth on the column that settles this contract and first on the agentic column it does not settle on. Our eight-model panel, run price-blind on August 26, had the two companies nearly level at 35 and 36.7 percent, and ten days of buying leave the OpenAI contract about ten points under that seat median rather than 20. The tape, the book behind every quote, and the four things that panel said would move this market are below.
Stay ahead of the markets.
Daily insights and expert picks on Kalshi, Polymarket, and what's moving markets.
Free forever. Unsubscribe anytime.
New to event markets? Our plain explanation of how prediction markets work covers the mechanics in about five minutes.
The One Column That Decides Everything
LiveBench scores every model it evaluates across seven categories, rewrites its questions periodically so nothing can be memorized, and publishes each category as its own column.
Kalshi's contract terms name exactly one of them. "The Underlying for this Contract is models on LiveBench.ai ranked by Coding Average score," the document reads, and the payout goes to whichever company "has the highest ranked model by Coding Average on [the date] at 10:00 AM ET."
Three details in that sentence do most of the work on this page, and the September 4 refresh turned the first one from a warning into a worked example.
Coding Average is only two tasks. LiveBench builds that column out of code generation and code completion. It publishes a completely separate column called Agentic Coding, built from three different tasks, and a Global Average across everything. GPT-6 Astra is the third-best model on the entire table by Global Average, at 83.05, behind only the two Anthropic Fable models. On the column that settles this contract it is fourteenth. A lab can ship the model the rest of the industry is calling the start of a new era and still lose this market, and on September 4 the market priced exactly that.
The snapshot is a moment, not a season. A model that leads for 11 months and gets passed on December 20 pays nothing, and the terms add that revisions made to the leaderboard after expiration "will not be accounted for."
A tie splits the dollar. If two listed companies are level at the top, the terms pay each side's Yes holders one dollar divided by the number tied, rounded down to the whole cent. On a table where third through seventh place are separated by under two points, that is a live outcome rather than a footnote.
Anthropic Still Holds First And Second, And The Top Four Have Not Moved
On the settlement view as refreshed September 10, the top of the Coding Average column reads: Claude Fable 5.1 Max Effort at 86.38, then Claude Fable 5 Max Effort at 85.99, then GPT-5.6 Sol at 83.94, then GPT-5.2 Codex at 83.62. That is the same four models in the same order as the table that caused the September 4 repricing.
When this page first published, Anthropic's lead over the best non-Anthropic model was 2.05 points and rested on one model. It is 2.44 now, and the company holds the top two slots and six of the top 15 — one fewer than on September 4, because the new DeepSeek row took the fifteenth place. For a contract that pays on one snapshot, the shape of that lead is what matters most: OpenAI's best model would have to pass two Anthropic models, not one, to take the column.
Anthropic announced Fable 5.1 and Mythos 5.1 on September 1. The company said Fable 5.1 "sets a new standard for coding, knowledge work, and long-running problem-solving tasks." It also said the two "are the same model, but with different levels of safeguards": Fable 5.1 is generally available, and Mythos 5.1 is offered only through the company's trusted-access programs.
There is a callback owed here. In August, four of our eight panel seats treated Anthropic's July safety-classifier redeployment of Fable 5 as a live threat to the leader's score, and one seat argued the dates foreclosed it. The table has now answered that debate more directly than any of them could: the successor scored 0.39 points above the model everyone was worried about. Our panel also listed, as one of the four things that would change its mind, "another Anthropic frontier model that scores above 85.99." This is the item that fired.
What the last ten days added was traffic, not evidence. Since this page last read the board, 4,952 Anthropic contracts have traded, and the aggressors split almost evenly: 2,733 lifting offers, 2,219 hitting bids. Prints ran as high as 72 and as low as 60. Open interest is down about 460 contracts from the 129,570 recorded on September 4, which is another way of saying those 4,952 contracts mostly changed hands between traders rather than creating net new exposure: opens and closes offset almost exactly. Its bid drifted three cents lower and its offer five, on a market that did not change its mind. A reader opening the Kalshi app tonight sees the opposite sign, and both readings are true. The exchange's own 24-hour reference reads 60 because 60 is where this contract sat at two o'clock on Sunday morning — Saturday traded as high as 71, bottomed at 64 and closed at 67, and the 60 is a single print 20 minutes into Sunday. Sunday was then spent buying it back: 317 contracts at 60 and 61 by late morning, and 13 minutes into Monday, 564 more from 67 up to 71. Measured against that 24-hour reference the contract is eight cents higher. Measured against this page's own last read it is three cents lower on the bid. The round trip is the story, and the sign depends entirely on where you start the clock.
OpenAI Is Ten Cents Higher Without A New Score
OpenAI still owns more of the neighborhood than any other company. It has four of the top seven scores, and GPT-5.2 Codex, a January model, is still fourth at 83.62 while posting only 49.39 on Agentic Coding — the profile of a model tuned narrowly for the exact two tasks this contract settles on.
GPT-6 Astra is the opposite profile, and it has not moved. Its 80.36 on Coding Average still puts it behind four of OpenAI's own older models: GPT-5.6 Sol, GPT-5.2 Codex, GPT-5.6 Luna and GPT-5.5 Thinking all score higher on the column. Fortune's launch report said the release was initially limited to a set of enterprise customers including OpenAI's cybersecurity-focused Daybreak program, with broader access, the API and AWS to follow "in the coming days," and 9to5Mac reported the same staged rollout into ChatGPT and Codex. Wider availability is not a score. LiveBench has not published a new OpenAI model since Astra, and there is still no Codex-branded model in the 5.6 or 6 generation anywhere on the table.
The historical pattern is unchanged and still the strongest argument OpenAI has. LiveBench has published 11 question sets since June 2024, and each one is its own scored table. Across all 11, the Coding Average lead has gone to Anthropic five times, to OpenAI five times, and to Google once. Nobody else, ever. Three of OpenAI's five turns at the top came from Codex-branded models. Seven of our eight panelists named that absence in August as the single most likely thing to flip the market, and ten days on it is still the case for the 27 cents.
What is new is who is holding the contract. Since this page's last read, 5,003 OpenAI contracts have traded and 3,739 of them, three in four, were buyers lifting offers. Open interest rose from 81,760 to 84,270. That combination is the tell: a contract climbing on net buying with open interest going up is being accumulated, not covered. Somebody is putting on a position rather than closing an old one. Nothing LiveBench published in those ten days gave them a reason.
The New Row: The Best Agentic Coder On The Table Sits Fifteenth On The One That Pays
The only model added inside the top 15 since September 4 is DeepSeek V4.1 Flash Max Effort. DeepSeek's own change log dates it September 10 and notes that "the previous-generation models V4 Flash and V4 Flash Vision Exp have been retired; for compatibility, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp are temporarily routed to V4.1 Flash." LiveBench's table carried it that same day.
It scores 80.04 on Coding Average, which is fifteenth, and moves DeepSeek's best model up 15 places from the V4 Pro that sat thirtieth on September 4. It scores 81.45 on Global Average, sixth on the table. And it scores 77.27 on Agentic Coding, the highest Agentic Coding score on the table — 11.21 points clear of Claude Fable 5.1, the model that leads the column this contract actually settles on.
It is not the only row that arrived. Grok 4.6's score has not changed, and it has slid from 31st on September 4 to 33rd — two places, which takes two new models above it, not one. DeepSeek's V4 Pro 0813 moved only from 30th to 31st over the same stretch, so the second arrival sits between them, and there is exactly one row there: Smaug Flash at 77.10, thirty-second, which LiveBench credits to Abacus.AI and lists as a fine-tune of DeepSeek V4 Flash 0731. Abacus.AI now holds three rows on this table, each of them someone else's weights tuned further and re-attributed to the tuner — the exact mechanism our August panel named as its fourth trigger, quietly demonstrated twice on one refresh.
That DeepSeek row is this page's thesis with a name attached. LiveBench's two coding columns rank the same 57 models differently enough that the best agentic coder on the board is the fifteenth-best coder by the only definition Kalshi's terms recognize. A trader who reads "best coding model" the way a developer reads it would buy the wrong contract, and the zero prints below say nobody has even made that mistake.
The market's response was nothing at all, and that is meant literally. DeepSeek's contract is 0 bid, 1 asked, and not one contract has traded on it since before the September 4 refresh — the row's best model climbed 15 places and its market did not print once. Our August panel gave DeepSeek 2.5 percent, which is 2.5 percent more than the market will pay for it.
The Tape: Fourteen Orders That Made The Ten Days
Below is every order of 450 contracts or more on this board since our September 4 read, plus the one that set OpenAI's high for the window. Times are Eastern; sizes are contracts of one dollar face value.
| Time (ET) | Contract | Prints | Contracts | Side |
|---|---|---|---|---|
| Sep 5, 1:34 P.m. | OpenAI | 21 | 659 | buyers |
| Sep 5, 9:13 P.m. | Anthropic | 66 to 67 | 778 | sellers |
| Sep 8, 10:54 A.m. | xAI | 7 | 500 | buyers |
| Sep 8, 2:45:48 P.m. | OpenAI | 23 | 602 | sellers |
| Sep 8, 2:45:53 P.m. | Anthropic | 66 | 500 | sellers |
| Sep 8, 7:08 P.m. | Anthropic | 66 to 68 | 684 | buyers |
| Sep 8, 7:55 P.m. | xAI | 7 to 9 | 604 | buyers |
| Sep 9, 2:41 P.m. | OpenAI | 23 | 627 | buyers |
| Sep 9, 4:29 P.m. | OpenAI | 26 to 27 | 530 | buyers |
| Sep 11, 5:46 P.m. | Anthropic | 63 to 67 | 488 | buyers |
| Sep 12, 10:28 P.m. | xAI | 7 to 29 | 1,413 | buyers |
| Sep 13, 9:46 A.m. | 2 | 1,268 | buyers | |
| Sep 13, 7:48 P.m. | OpenAI | 27 to 34 | 293 | buyers |
| Sep 14, 12:13 A.m. | Anthropic | 67 to 71 | 564 | buyers |
Each row is one order: every fill in it shares the same millisecond timestamp and the same aggressor side. "Buyers" means the order lifted resting offers; "sellers" means it hit resting bids. A row with a price range is that order taking more than one price level.
Look twice at September 12 at 10:28 p.m., when a single xAI buyer took 1,413 contracts in one second and paid as much as 29 cents for a contract that had traded between 6 and 9 all week. By the next morning xAI was printing 6 again, and it has not printed above 6 since. Nothing was announced about xAI that night, and Grok 4.6 has not moved on the table — its 76.78 is unchanged, and it sits 33rd now rather than 31st only because other models were scored above it. What that order found was an empty offer stack, and the book in the next section shows exactly how empty.
The September 13 row underneath it is the one with a real change behind it. Google's contract had not traded since September 2 — ten straight Eastern-time sessions, September 3 through September 12, with zero volume and no bid posted at any price. Then at 9:46 a.m. one buyer took 1,268 contracts at 2 cents, another took 312 at 3 that afternoon, and open interest rose 1,616. Ten days ago Google was the emptiest row on this board. Tonight it carries a resting bid for 2,756 contracts at 2 cents, where before there was no bid at any price.
Worth one line for the pair five seconds apart on September 8: a 602-contract seller in OpenAI at 2:45:48 p.m. and a 500-contract seller in Anthropic at 2:45:53. Two different contracts, five seconds apart, sold in the same direction — the only time in the ten days that two orders this size hit two different contracts inside the same minute.
What The Two-Sided Quotes Actually Say
Now the promise from the top of the page, twice over.
Anthropic's quote is honest. The 66 bid holds 99 contracts, and behind it sit 100 at 65, 102 at 61, 123 at 60, 153 at 59 and 250 at 58. The 68 offer holds 43, then 34 at 70 and a wall of 461 at 71. A buyer who wants five hundred contracts pays up to 71; a seller who wants five hundred takes 59 for the last of them. The quoted market is 66/68 and the real market for size is roughly 59 to 71 — wide, but two-sided the whole way down.
OpenAI's quote is not honest, and this is the number worth carrying away from the page. The contract is 27 bid, 28 asked. There are 27 contracts bid at 27. The 28 offer holds one contract. Behind that single contract the next offers are 85 at 31, 26 at 32, 101 at 34 and 400 at 35. Behind the 27 bid: 15 at 26, 25 at 24, 23 at 23, 35 at 21, 64 at 20, then 500 at 19 and 497 at 18.
Compare that with what this page recorded on September 4, when OpenAI was 17 bid, 18 asked with 627 contracts bid and 232 offered. That was also a one-cent market, and it was a real one. The price is ten cents higher tonight and the depth at it has fallen by two orders of magnitude. For anything above a token trade the true market is closer to 19 bid, 31 offered — a twelve-point spread wearing a one-cent mask. The September 13 order that printed 34 is the proof: it needed only 293 contracts to move the price seven cents, and the price was back at 27 within a day.
The book, in one lineOpenAI is quoted 27 bid, 28 asked, and that 28-cent offer holds one contract. The five hundredth contract costs 35.
A Worked Example: What 500 Contracts Actually Cost Tonight
The quote is what one contract costs. Here is what five hundred costs, walked straight up each ask ladder above.
| Contract | Quoted ask | Levels a 500-lot buyer clears | All-in cost | Average price paid |
|---|---|---|---|---|
| Anthropic | 68¢ | 43 at 68, 33.76 at 70, 423.24 at 71 | $353.37 | 70.7¢ |
| OpenAI | 28¢ | 1 at 28, 85 at 31, 26 at 32, 101.36 at 34, 286.64 at 35 | $169.74 | 33.9¢ |
Anthropic's average fill lands 2.7 cents above its quote. OpenAI's lands 5.9 above, better than a fifth on top of the screen price, and the five hundredth OpenAI contract is bought seven cents higher than the first. Call it the difference between a thin book and a hollow one. It is also why the same "one-cent market" label means two completely different things on the two contracts that matter here.
The same emptiness sits under xAI, which is 6 bid, 7 asked with 428 contracts offered at 7 and then nothing at all until 154 contracts at 20. Into that gap fell the 1,413-contract order of September 12. On the bid side xAI carries a resting order to buy 10,104 contracts at a penny: $101 of actual money, and it can be pulled the moment it is needed. Call it optionality, not a floor.
Read across the whole board, the nine asks now sum to 113 and the nine bids to 103. Both of this page's earlier readings showed a 13-point width between those sums — 110/97 in August, 107/94 on September 4. It is 10 tonight. The board is carrying more implied probability than it was and quoting it more tightly, which is what happens when a row that had no bid at all grows one.
The Board, Marked To Market
The board as of 2:00 a.m. ET on September 14, 2026, alongside where it stood at our last read. Market prices are the bid and the ask from the live order book. The panel column is the August 26 seat median after the revision round, run price-blind and left unchanged so it can be graded honestly against what happens; with eight seats the median is the midpoint of the two middle answers, which is what keeps one outlier from dragging a blended average around. Leaderboard positions are from the September 10 table.
| Company | Best model on the settlement column | Market, Sept. 14 (bid / ask) | Market, Sept. 4 (bid / ask) | Panel, Aug. 26 |
|---|---|---|---|---|
| Anthropic | Claude Fable 5.1, 1st | 66¢ / 68¢ | 69¢ / 73¢ | 35.0% |
| OpenAI | GPT-5.6 Sol, 3rd | 27¢ / 28¢ | 17¢ / 18¢ | 36.7% |
| xAI | Grok 4.6, 33rd | 6¢ / 7¢ | 6¢ / 7¢ | 1.1% |
| Gemini 3.7 Flash, 21st | 2¢ / 3¢ | 0¢ / 2¢ | 6.8% | |
| Z.ai | GLM-5.2, 17th | 1¢ / 2¢ | 1¢ / 2¢ | 2.0% |
| Moonshot AI | Kimi K3, 10th | 1¢ / 2¢ | 1¢ / 2¢ | 5.0% |
| DeepSeek | DeepSeek V4.1 Flash, 15th | 0¢ / 1¢ | 0¢ / 1¢ | 2.5% |
| Alibaba | Qwen 3.6 Plus, 26th | 0¢ / 1¢ | 0¢ / 1¢ | 1.8% |
| Baidu | none on the leaderboard | 0¢ / 1¢ | 0¢ / 1¢ | 0.2% |
| No Listed Company | an unlisted organization wins | not tradeable | not tradeable | 9.5% |
Six of the nine contracts have a real quote on both sides. DeepSeek, Alibaba and Baidu show no bid at all, so their one-cent ask is a displayed price rather than a two-sided market with someone waiting on the other end. Kalshi reports volume and open interest in dollars of contract face value, one dollar per contract, the way the exchange itself displays them.
Two rows deserve narration, and they point in opposite directions. The first is OpenAI, where the frozen panel's 36.7 percent has not moved since August 26 and the market has spent ten days walking ten cents back toward it — from about 20 points below the panel to about ten. Nobody re-argued anything; the market simply decided its own reaction to the September 4 table had gone too far. The second is DeepSeek, the row with the only genuinely new fact on this page. Its best model climbed 15 places on the settlement column and took first on the agentic one, and its quote did not change by a penny. Google's row is the one I keep coming back to: the panel's 6.8 percent rests entirely on a Gemini Pro line that TechCrunch reported in July had last been updated in February, and no Pro has arrived since — yet Google is the row that woke up.
The unlisted field is unchanged and still the quietest large number on the board. Smaug Agentic, a fine-tune of Moonshot's Kimi K3 that LiveBench credits to Abacus.AI, is sixth at 82.47 — fifth when the panel priced it in August. Meta's Muse Spark 1.3 is twelfth at 81.06. Both sit above GPT-6 Astra. If either organization holds the top score at the snapshot, all nine contracts pay zero. The panel gave that 9.5 percent in August; the market still cannot sell it to you.
Every seat's number, before and after the August 26 revision round. The row that travelled furthest is the DeepSeek seat, which came into the second round ten points low on Anthropic and 15 points high on the unlisted field, and left it inside the pack on both:
| Seat | Anthropic | OpenAI | Moonshot | No listed co. | |
|---|---|---|---|---|---|
| Claude Fable 5 | 40.0 → 36.0 | 36.0 → 33.0 | 8.2 → 6.5 | 3.5 → 5.0 | 5.0 → 9.0 |
| Claude Opus 5 | 40.0 → 39.5 | 35.0 → 36.3 | 7.0 → 6.5 | 4.0 → 4.0 | 7.5 → 7.0 |
| Claude Sonnet 5 | 37.0 → 35.0 | 35.0 → 37.0 | 5.0 → 6.0 | 6.0 → 4.5 | 6.5 → 9.0 |
| ChatGPT | 36.0 → 36.5 | 32.0 → 35.0 | 9.0 → 7.0 | 6.0 → 5.0 | 6.5 → 9.0 |
| Gemini | 34.0 → 34.0 | 44.0 → 38.0 | 2.0 → 5.0 | 3.0 → 6.0 | 8.0 → 10.0 |
| GLM | 28.0 → 33.0 | 32.0 → 36.0 | 4.0 → 7.0 | 8.0 → 5.0 | 17.5 → 12.0 |
| Kimi | 36.0 → 33.0 | 32.0 → 37.0 | 7.0 → 7.0 | 5.0 → 5.0 | 11.5 → 10.0 |
| DeepSeek | 25.0 → 35.0 | 35.0 → 38.0 | 8.0 → 7.0 | 3.0 → 5.0 | 25.0 → 10.0 |
| Panel Middle | 36.0 → 35.0 | 35.0 → 36.7 | 7.0 → 6.8 | 4.5 → 5.0 | 7.7 → 9.5 |
Seat numbers are each model's own percentages as submitted on August 26. One first-round board came back summing to 100.2 rather than 100, so every seat is normalized to sum to 100 before the middle is taken, which is why the bottom row is not always the plain midpoint of the column above it, and the medians are taken column by column, so they do not sum to exactly 100. The seats were not re-run for this update; the Claude Fable 5 seat is the model the current table lists second, behind its own successor.
Records to date for the four externally graded seats on this panel, each graded against real market settlements: Gemini 86% on 8,779 graded calls · GLM 83% on 4,020 graded calls · Kimi 86% on 4,394 graded calls · DeepSeek 84% on 4,083 graded calls. Recomputed daily; the full scoreboard is public.
These are model estimates, not predictions of fact and not financial advice. Kalshi is a CFTC-regulated exchange for event contracts, 18+ only, and availability varies. Every number in this piece gets graded in public once the market settles: the running record lives on the full graded scoreboard.
More live boards from the same panel: the broader best AI at the end of 2026 board, where the market is not restricted to coding · the top-ranked AI model board, where the settlement leaderboard is a different one entirely · the GPT-6 release-date ladder, which the September 3 launch answered · the Anthropic IPO ladder, where the before-November-1 rung has fallen to 55¢ bid · and Kalshi's daily prediction-market hub for what the panel priced this morning.
Hottest Prediction Markets Right Now
- 2028 Democratic presidential nominee · $194M traded
- 2027 Pro Football Champion · $72M traded
- 2028 U.S. Presidential Election winner? · $62M traded
- Pro Baseball Champion · $60M traded
- 2028 Republican presidential nominee · $60M traded
- Bitcoin price at the end of 2026 · $33M traded
Every market above links to our full AI model verdict; browse them all on the OddsShopper prediction markets hub, and see how every settled call actually scored on the full graded scoreboard.
The IPO Ladder Answered The Question This Page Asked
On September 4 this page looked at Kalshi's other Anthropic board — the ladder on when the company's IPO is confirmed — and argued with one rung of it. "An 82-cent price says the calendar lands on the near side of it," it said of the before-November-1 contract. "Reuters' wording does not say that."
That rung is 55 bid, 56 asked tonight. December 1 has come down from 86 bid to 71/73, and October 1, which was 3 bid on September 4 after more than 45,000 contracts traded in a day, is 1/3. The three rungs Kalshi listed that week for October 10, 17 and 24 — "quoted too wide to read" when we last looked — now carry two-sided markets at 3/5, 12/13 and 34/36, a proper curve where there was a blank.
The reporting behind those prices did not change. Reuters' September 5 story, which CNBC also carried, remains the operative account: marketing in mid-October at the earliest, a listing days before the midterms, and a public prospectus "not expected until late September." What changed is where the size went. The before-November-1 rung turned over 19,822 contracts in 24 hours against 83,103 open, and the before-October-1 rung 23,123 against 328,940 — the two rungs our argument was about are the two carrying the volume, and that single November rung traded roughly six times the whole nine-contract coding board's 24-hour total of 3,217. Read across the midpoints, the back half of October is now the modal window: the week ending October 24 carries 22 points of the curve and the week ending November 1 another 21, close enough that the spread cannot separate them. That is what the calendar in the Reuters paragraph implies, and it is not what an 82-cent November rung said.
And because this page argues that depth explains moves, here is that rung's depth, which is the mirror image of the coding board's. The 55 bid holds 230 contracts, with 1,290 behind it at 54, 1,045 at 52 and 1,000 at 50. Above the quote sit 1,284 contracts at 56, then 465 at 58, 2,109 at 59 and 4,800 at 60. A buyer of a thousand contracts there is done by 56. On the OpenAI coding contract, a thousand contracts would clear every offer up to 51 cents against a quoted ask of 28. Same exchange, same night, same company somewhere in both questions — and completely different markets underneath the quotes.
Where The Panel Changed Its Mind
The revision round moved real numbers on this board, and no money changed hands doing it. The clearest example came from the seat that started furthest from everyone else. Where the seats say FIELD below, they mean the last row of the board: nobody with a contract holds the top score, so every one of the nine pays zero.
"The FIELD deserves high probability (25%) because unlisted organizations like Abacus.AI (82.47) are highly competitive, and the board excludes major players like Meta." — DeepSeek, first round
After reading the other seven arguments, it cut that to 10 and rewrote the reasoning around a narrower mechanism, while moving Anthropic up ten points:
"FIELD deserves 10% because Abacus.AI's fine-tune already ranks fifth, demonstrating unlisted organizations can compete for this narrow column." — DeepSeek, second round
Read that first-round quote again with the current table in hand. The seat named Meta as the excluded player to worry about, and Meta's Muse Spark 1.3 was on the next table LiveBench posted, in twelfth.
The most interesting revision was an argument nobody else had made. Four of the eight seats treated Anthropic's July 1 safety-classifier redeployment as a live threat to the leader's score. One seat went back and checked the dates:
"A, C, D and H all treated Anthropic's July 1 classifier redeploy as a live downside on the 85.99. But the timeline in the card forecloses that: Fable 5 was suspended June 12, the question set is dated June 25, and site data was refreshed August 25, so 85.99 is almost certainly the post-classifier score already." — Claude Opus 5, second round
[Editor's note: LiveBench does not publish a run date beside any individual model score, so this could not be confirmed from the settlement source in August. The September 4 table settled the practical question a different way: the successor model scored 86.38, above the 85.99 the seats were arguing about.]
And one seat moved with the panel, from 5 to 9 on the unlisted field, but for a reason nobody else gave. Its argument turns on a loophole in who gets credit: anyone can take a strong open-weight model, tune it further, and be listed as the owner of the result.
"I hold FIELD at 9, above most of the panel. Abacus.AI already sits fifth at 82.47 by fine-tuning Moonshot's open weights, LiveBench attributes fine-tunes to the tuner, and any December open-weight frontier drop invites the same arbitrage." — Claude Fable 5, second round
What The Panel Said Would Move It, And What Has
Four things, each of them checkable rather than atmospheric. One fired on September 4. The one the panel rated most likely still has not, and the row that moved most in the last ten days was not on the list at all.
A Codex-branded release from OpenAI. Seven of eight seats named this as the highest-probability crown flip in the window. It has not happened. GPT-6 Astra is not Codex-branded, it scored 80.36 on the column, and ten days later there is still no Codex in the 5.6 or 6 generation on LiveBench's table — only January's GPT-5.2 Codex, fourth at 83.62. A Codex variant of the 6 generation, scored above 86.38, remains the most likely OpenAI path the table shows, and the buyers who took the contract from 17 to 27 are paying ahead of it rather than after it. Direction if it lands: pushes OpenAI up and everything else down.
A Gemini Pro release before December. Google's price and its leaderboard position are both explained by a Pro line that, as TechCrunch reported on July 21, had last been updated in February, with Bloomberg reporting internal delays on a 3.5 Pro. Nothing has changed on the table — Gemini 3.1 Pro Preview is still Google's best Pro entry, 34th at 76.45, and the company's best model of any kind is the 3.7 Flash at 78.89, 21st. The contract, however, has changed: no trade and no bid for ten sessions, then 1,629 contracts on September 13 with 1,580 of them buyer-side, and a 2,756-contract bid now standing at 2 cents. Somebody is paying for the possibility on the strength of no announcement whatsoever. Direction if it lands: pushes Google up sharply, Anthropic and OpenAI down.
Another Anthropic frontier model that scores above 85.99. This is the one that fired. Fable 5.1 scored 86.38 on September 4 and Anthropic repriced from 50 to 66 within 15 minutes. Since then the score has held and the quote has eased from 69/73 to 66/68, which is the market's way of saying it believes the table and is less sure about the four months between now and the snapshot. The ceiling question is unchanged: whether 66 to 68 cents already prices a hold through December 31.
An open-weight frontier release in the fourth quarter. This leg has not fired either, but its mechanism got a fresh demonstration. LiveBench's own registry credits three Abacus.AI entries — Smaug Agentic at 82.47, Smaug Flash at 77.10 and Smaug Mini at 75.43 — to Abacus.AI rather than to the labs whose weights they are built on, which the registry lists as Kimi K3, DeepSeek V4 Flash 0731 and Qwen3.8-27B. The best of the three sits sixth on the settlement column, above GPT-6 Astra. The attribution rule the payout depends on is the tuner's, not the base lab's, and the September 10 refresh added a fresh instance of it: Smaug Flash, thirty-second, tuned from a DeepSeek Flash checkpoint. Direction: pushes the unlisted field up and every listed company down.
Settlement timeline
| Date | What happens |
|---|---|
| June 25, 2026 | The question set the contract currently settles against |
| September 1, 2026 | Anthropic releases Claude Fable 5.1 and Mythos 5.1 |
| September 3, 2026 | OpenAI releases GPT-6 Astra |
| September 4, 2026, 4:09 P.m. ET | LiveBench refreshes: Fable 5.1 first at 86.38, GPT-6 Astra fourteenth at 80.36 |
| September 10, 2026 | DeepSeek releases V4.1 Flash; LiveBench's table carries it the same day, fifteenth on Coding Average and first on Agentic Coding |
| September 13, 2026 | Google's contract trades for the first time since September 2: 1,629 contracts, 97 percent of them buyer-side, at 2 and 3 cents |
| December 31, 2026, 10:00 A.m. ET | The snapshot; last trading time is the same moment |
| No Later Than January 1, 2027 | Settlement, unless the outcome goes to review |
September 10 is the row in that table with the shortest fuse: the settlement source added a model the same day it shipped, and nobody had to announce anything to this market for it to happen. This page is re-scored when the story moves. Market prices on it are stamped September 14, 2026; panel numbers are stamped August 26, 2026.
More on this: Where Does Bitcoin End 2026? Kalshi's Price-Band Odds · Nurmagomedov Vs Song Odds: Not SO Fast On Umar · Kalshi Venezuela Leader Odds: Who Holds Power In 2026? · Why Andy Barr (R) Will Win the Kentucky Senate Race · Why Jeff Merkley (D) Will Win the Oregon Senate Race
The Bottom Line
For ten days this market has had nothing new at the top of its leaderboard to price, and it has repriced anyway. The top four scores on the Coding Average column are the same four in the same order as the table that caused the September 4 crash. OpenAI has not been scored higher since. And yet OpenAI's quote is ten cents up, Anthropic's is three to five cents down, Google has gone from no bid at any price to 2,756 contracts bid at 2 cents, and an xAI contract that traded between 6 and 9 all week printed 29 in a single second on a Saturday night.
Every one of those moves came out of the book rather than the benchmark, which is why the depth matters more than the quote here. The 28-cent OpenAI offer holds one contract. The 71-cent Anthropic offer holds 461. Two contracts on the same board, both quoted to the penny, and only one of them will let you act on what you think.
The genuinely new fact in the window belongs to the row nobody priced. DeepSeek shipped a model on September 10 that is the best agentic coder on this question set's table and the fifteenth-best coder by the definition this contract uses, and it is still bid zero. Between now and December 31 there will be more refreshes than headlines, and the refreshes are the only ones that pay. The last one put the best agentic coder on the table into a seat that pays nothing — which is exactly the reading error this market exists to charge people for.



