Updated September 4, 2026 · 15 min read · by Sam Smith. First published August 26. Panel verdicts were generated August 26, price-blind, from the LiveBench.ai settlement view and the contract terms as published by the exchange; they have not been re-run and are shown here as the frozen read. Kalshi prices, order book and trade tape fetched September 4, 2026, 10:17 p.m. ET (September 5, 02:17 UTC). LiveBench table as refreshed September 4 at 4:09 p.m. ET.
OpenAI launched GPT-6 Astra on Thursday afternoon, and its president, Greg Brockman, told reporters that "it's not unreasonable to feel that we are now in the AGI era," as Fortune reported. Traders on Kalshi's best-coding-model board had been buying OpenAI into that launch all week. The contract that printed 29 cents on Tuesday evening printed 40 cents at 2:13 p.m. Friday.
At 4:09 p.m. Friday, LiveBench refreshed the table this market settles on. The new top score on the Coding Average column belonged to Claude Fable 5.1 Max Effort at 86.38, a model Anthropic had announced on Tuesday. GPT-6 Astra Max Effort was on the table too. It scored 80.36 on that column, fourteenth, behind four of OpenAI's own older models and behind a fine-tune from a company with no contract on this board.
Ten minutes later the first OpenAI print after the refresh was a seller, and every OpenAI print for the rest of the night was a seller. By 4:33 p.m. one order had pushed more than 4,000 contracts through the bids from 25 cents down to 21. The evening's last prints were 14 and 15. Anthropic went the other way in the same quarter hour, from 52 to 66, and touched 75 before the night was over.
This is a market that trades a leaderboard refresh, not a product launch, and it rewards the trader who reads the column over the one who reads the headline. The open question after Friday is whether 69 to 73 cents is a correct repricing of a two-model Anthropic lead or an overreaction on a thin book. The tape and the order book below are the evidence for each side, and the size of the move and the size of the book that made it are very different numbers.
The Quick Answer
Anthropic is quoted 69 cents bid, 73 cents ask on Kalshi's board for the best AI coding model at the end of 2026, up from 57 to 60 when this page first published on August 26. OpenAI is 17 bid, 18 ask, down from 26 to 29. The move happened in about 25 minutes on Friday afternoon, immediately after LiveBench posted Claude Fable 5.1 as the new leader of its Coding Average column and GPT-6 Astra in fourteenth. The 73-cent offer holds 15 contracts, so the price for size is 75. Our eight-model panel, run price-blind on August 26, had the two companies nearly level at 35 and 37 percent. The full tape, the order book behind the quotes, the refreshed leaderboard and the Anthropic IPO ladder that moved the same day are below.
Stay ahead of the markets.
Daily insights and expert picks on Kalshi, Polymarket, and what's moving markets.
Free forever. Unsubscribe anytime.
New to event markets? Our plain explanation of how prediction markets work covers the mechanics in about five minutes.
The One Column That Decides Everything
LiveBench scores every model it evaluates across seven categories, rewrites its questions periodically so nothing can be memorized, and publishes each category as its own column.
Kalshi's contract terms name exactly one of them. "The Underlying for this Contract is models on LiveBench.ai ranked by Coding Average score," the document reads, and the payout goes to whichever company "has the highest ranked model by Coding Average on [the date] at 10:00 AM ET."
Three details in that sentence do most of the work on this page, and Friday's refresh turned the first one from a warning into a worked example.
Coding Average is only two tasks. LiveBench builds that column out of code generation and code completion. It publishes a completely separate column called Agentic Coding, built from three different tasks, and a Global Average across everything. GPT-6 Astra is the third-best model on the entire table by Global Average, at 83.05, behind only the two Anthropic Fable models. On the column that settles this contract it is fourteenth. A lab can ship the model the rest of the industry is calling the start of a new era and still lose this market, and on Friday the market priced exactly that.
The snapshot is a moment, not a season. A model that leads for 11 months and gets passed on December 20 pays nothing, and the terms add that revisions made to the leaderboard after expiration "will not be accounted for."
A tie splits the dollar. If two listed companies are level at the top, the terms pay each side's Yes holders one dollar divided by the number tied, rounded down to the whole cent. On a table where third through seventh place are separated by under two points, that is a live outcome rather than a footnote.
Fable 5.1 Debuted On Top, And Anthropic Now Holds First And Second
On the settlement view as refreshed Friday afternoon, the top of the Coding Average column reads: Claude Fable 5.1 Max Effort at 86.38, then Claude Fable 5 Max Effort at 85.99, then GPT-5.6 Sol at 83.94, then GPT-5.2 Codex at 83.62.
When this page first published, Anthropic's lead over the best non-Anthropic model was 2.05 points and rested on one model. It is now 2.44 points, and the company holds the top two slots, seven of the top 15 overall. The crown no longer rests on a single score. For a contract that pays on one snapshot, this is the change that matters most, because it means OpenAI's best model would now have to pass two Anthropic models, not one, to take the column.
Anthropic announced Fable 5.1 and Mythos 5.1 on September 1. The company said Fable 5.1 "sets a new standard for coding, knowledge work, and long-running problem-solving tasks." It also said the two "are the same model, but with different levels of safeguards": Fable 5.1 is generally available, and Mythos 5.1 is offered only through the company's trusted-access programs. Fable 5.1 did not exist on the LiveBench table when we pulled it on August 25. It was there on Friday, at the top.
There is a callback owed here. In August, four of our eight panel seats treated Anthropic's July safety-classifier redeployment of Fable 5 as a live threat to the leader's score, and one seat argued the dates foreclosed it. The table has now answered that debate more directly than any of them could: the successor scored 0.39 points above the model everyone was worried about. Our panel also listed, as one of the four things that would change its mind, "another Anthropic frontier model that scores above 85.99." That is the item that fired.
OpenAI Shipped GPT-6 Astra, And It Landed Fourteenth On The Column That Pays
OpenAI still owns more of the neighborhood than any other company. It has four of the top seven scores, and GPT-5.2 Codex, a January model, is still fourth at 83.62 while posting only 49.39 on Agentic Coding. That profile describes a model tuned narrowly for the exact two tasks this contract settles on.
GPT-6 Astra is the opposite profile. Its 57.32 on Agentic Coding and 83.05 on Global Average are the numbers of a broad frontier model. Its 80.36 on Coding Average puts it behind four of OpenAI's own older models. GPT-5.6 Sol, GPT-5.2 Codex, GPT-5.6 Luna and GPT-5.5 Thinking all score higher on the column. Fortune's launch report said the release was initially limited to a set of enterprise customers including OpenAI's cybersecurity-focused Daybreak program, with broader access, the API and AWS to follow "in the coming days," so a Codex-branded sibling could still be scored later. For now, the model OpenAI built to reset the frontier is the eleventh-best coding-column entry from the two companies that matter here.
The historical pattern is unchanged and still the strongest argument OpenAI has. LiveBench has published 11 question sets since June 2024, and each one is its own scored table. Across all 11, the Coding Average lead has gone to Anthropic five times, to OpenAI five times, and to Google once. Nobody else, ever. Three of OpenAI's five turns at the top came from Codex-branded models, and there is still no Codex in the 5.6 or 6 generation on the table. Seven of our eight panelists named that absence in August as the single most likely thing to flip the market. It remains the case for the 17 cents.
The Tape: Ten Minutes From Leaderboard To Repricing
Here is Friday afternoon on the two contracts that matter, in Eastern time, with sizes in contracts of one dollar face value. LiveBench's data stamp on the refreshed table reads 4:09:45 p.m. The last OpenAI trade before that stamp was at 2:13 p.m., at 40 cents; the last Anthropic trade before it was at 3:42 p.m., at 50.
| Time (ET) | Contract | Prints | Contracts | Side |
|---|---|---|---|---|
| 4:19:56 P.m. | OpenAI | 38 down to 34 | 1,541 | sellers |
| 4:20:00 P.m. | Anthropic | 52 | 204 | buyers |
| 4:20:40 P.m. | OpenAI | 33 to 32 | 508 | sellers |
| 4:23:04 P.m. | Anthropic | 54 up to 63 | 699 | buyers |
| 4:24:03 P.m. | Anthropic | 63 up to 66 | 278 | buyers |
| 4:24:11 P.m. | OpenAI | 30 to 28 | 675 | sellers |
| 4:24:59 P.m. | OpenAI | 28 to 27 | 536 | sellers |
| 4:33:54 P.m. | OpenAI | 25 down to 21 | 4,047 | sellers |
| 5:16 P.m. | OpenAI | 16 | 12 | sellers |
| 6:30 P.m. | OpenAI | 15 to 14 | 52 | sellers |
| 8:22 P.m. | Anthropic | 70 up to 75 | 33 | buyers |
| 9:00 P.m. | Anthropic | 75 | 32 | buyers |
| 10:11 P.m. | Anthropic | 70 to 69 | 8 | sellers |
Each row with a price range is a single order walking the book: every fill in it shares the same millisecond timestamp and the same aggressor side. "Sellers" means the order hit resting bids; "buyers" means it lifted resting offers.
The row that tells the story is the one at 4:33:54 p.m. One seller put 4,047 contracts through the OpenAI bids in a single sweep, starting at 25 and finishing at 21, and 2,500 of them filled at the bottom. From the leaderboard stamp to the close of that sweep is 24 minutes. Across the whole evening after the refresh, 7,371 OpenAI contracts traded and every one of them was a seller hitting a bid; not a single buyer crossed the spread. Anthropic's post-refresh tape is the mirror image: 1,255 contracts, 1,247 of them buyers, until eight contracts of profit-taking at 10:11 p.m. set the closing prints at 70 and 69.
The week before the refresh matters because it is what got sold. OpenAI's climb from 29 to 40 was built on buyers. On Tuesday 1,634 contracts traded, all of them lifting offers, including a single 1,008-contract order at 29 at 6:37 p.m. Wednesday added 2,212 contracts, 2,042 of them buyers, and a 35 close. Thursday printed 38, and Friday afternoon printed the 40. Anthropic had its own scare on Thursday night, when a 955-contract sell sweep at 8:24 p.m. pushed it from 55 to 41 in 17 seconds. That 41 is what Kalshi's day-over-day reference price still reads for Anthropic, and 37 for OpenAI, which is why a move measured from those references to tonight's bids looks like 28 points and 20 points. Those references are single prints. The honest measure is the quotes: Anthropic from 57/60 when this page published to 69/73; OpenAI from 26/29 to 17/18.
What The Two-Sided Quotes Actually Say
Now the promised gap. Anthropic's open interest on August 26 was 128,472 contracts. On Friday night it was 129,570, a net change of about 1,100. Every Anthropic trade since this page first published adds up to 5,000 contracts, and the entire post-refresh repricing from the 50 pre-refresh print to the 69 close took 1,255 of them, under one percent of the open interest. Measured since this page first published, the quote moved 12 points on turnover equal to about four percent of the money already in the contract. OpenAI turned over more, 15,983 contracts since August 26 against 81,760 open, but its repricing was still done by a handful of sellers in one afternoon.
The order book behind the 69/73 quote explains how a market that size can move that far on that little. The 69 bid holds 93 contracts. Behind it are 161 at 68 and 500 at 67. The 73 offer holds 15 contracts, and the next real offer is 729 contracts at 75. So the price a buyer actually pays for size tonight is 75, not 73, and the price a seller actually gets for size is closer to 67. The quoted midpoint is real; the depth at it is not.
The book, in one line: the 73-cent offer holds 15 contracts. Anyone buying size pays 75.
OpenAI's book is deeper on the bid than the price suggests. There are 627 contracts bid at 17 and 232 offered at 18, a one-cent market. Below that, 100 at 16 and, past a few small bids, a single resting order to buy 10,000 contracts at 10 cents, more contracts than the entire post-refresh selling but only $1,000 of actual money, and it can be pulled the moment it is needed. Call it optionality, not a floor.
Read across the whole board, the nine asks now sum to 107, down from 110 in August, but the nine bids fell the same three points, to 94 from 97. The total width is unchanged at 13. The overround moved; the market did not get tighter.
The Board, Marked To Market
Here is the board as of 10:17 p.m. ET on Friday, September 4, 2026, alongside where it stood when this page published. Market prices are the bid and the ask from the live order book. The panel column is the August 26 seat median after the revision round, run price-blind and left unchanged so it can be graded honestly against what happened; with eight seats the median is the midpoint of the two middle answers, which is what keeps one outlier from dragging a blended average around. Leaderboard positions are from Friday's refreshed table.
| Company | Best model on the settlement column | Market, Sept. 4 (bid / ask) | Market, Aug. 26 (bid / ask) | Panel, Aug. 26 |
|---|---|---|---|---|
| Anthropic | Claude Fable 5.1, 1st | 69¢ / 73¢ | 57¢ / 60¢ | 35.0% |
| OpenAI | GPT-5.6 Sol, 3rd | 17¢ / 18¢ | 26¢ / 29¢ | 36.7% |
| xAI | Grok 4.6, 31st | 6¢ / 7¢ | 10¢ / 11¢ | 1.1% |
| Z.ai | GLM-5.2, 16th | 1¢ / 2¢ | 2¢ / 3¢ | 2.0% |
| Moonshot AI | Kimi K3, 10th | 1¢ / 2¢ | 1¢ / 2¢ | 5.0% |
| Gemini 3.7 Flash, 20th | 0¢ / 2¢ | 1¢ / 2¢ | 6.8% | |
| DeepSeek | DeepSeek V4 Pro, 30th | 0¢ / 1¢ | 0¢ / 1¢ | 2.5% |
| Alibaba | Qwen 3.6 Plus, 25th | 0¢ / 1¢ | 0¢ / 1¢ | 1.8% |
| Baidu | none on the leaderboard | 0¢ / 1¢ | 0¢ / 1¢ | 0.2% |
| No Listed Company | an unlisted organization wins | not tradeable | not tradeable | 9.5% |
Five of the nine contracts have a real quote on both sides. Google, DeepSeek, Alibaba and Baidu show no bid at all, so their one- or two-cent ask is a displayed price rather than a two-sided market with someone waiting on the other end. Kalshi reports volume and open interest in dollars of contract face value, one dollar per contract, the way the exchange itself displays them.
Two rows deserve narration. The first is the gap between the market and the frozen panel on Anthropic, which has gone from 22 points to 34. Friday's leaderboard is evidence for the market's side of this one, and the seat records below are how that gets graded in public. The second is xAI, which has fallen from third to third-by-default: its best model, Grok 4.6 at 76.78, dropped from 27th to 31st as new models joined the table above it, and it still trades at 6 cents while Google, which has actually held this crown, has no bid at all. Google's row is the one I keep coming back to: the panel's 6.8 percent rests entirely on a Gemini Pro line that TechCrunch reported in July had last been updated in February, and Friday's table added a Gemini 3.8 Flash but no Pro.
The unlisted field got bigger, and it is worth a sentence of its own. Fifth place in August belonged to Smaug-Agentic at 82.47, a fine-tune of Moonshot's Kimi K3 that LiveBench credits to Abacus.AI. It is sixth now, and Meta's Muse Spark 1.3 arrived on Friday's table in twelfth at 81.06, above GPT-6 Astra. Two organizations with no contract on this board now sit inside the top 12 of the column that decides it. If either holds the top score at the snapshot, all nine contracts pay zero. The panel gave that 9.5 percent in August; the market still cannot sell it to you.
Every seat's number, before and after the August 26 revision round:
| Seat | Anthropic | OpenAI | Moonshot | No listed co. | |
|---|---|---|---|---|---|
| Claude Fable 5 | 40.0 → 36.0 | 36.0 → 33.0 | 8.2 → 6.5 | 3.5 → 5.0 | 5.0 → 9.0 |
| Claude Opus 5 | 40.0 → 39.5 | 35.0 → 36.3 | 7.0 → 6.5 | 4.0 → 4.0 | 7.5 → 7.0 |
| Claude Sonnet 5 | 37.0 → 35.0 | 35.0 → 37.0 | 5.0 → 6.0 | 6.0 → 4.5 | 6.5 → 9.0 |
| ChatGPT | 36.0 → 36.5 | 32.0 → 35.0 | 9.0 → 7.0 | 6.0 → 5.0 | 6.5 → 9.0 |
| Gemini | 34.0 → 34.0 | 44.0 → 38.0 | 2.0 → 5.0 | 3.0 → 6.0 | 8.0 → 10.0 |
| GLM | 28.0 → 33.0 | 32.0 → 36.0 | 4.0 → 7.0 | 8.0 → 5.0 | 17.5 → 12.0 |
| Kimi | 36.0 → 33.0 | 32.0 → 37.0 | 7.0 → 7.0 | 5.0 → 5.0 | 11.5 → 10.0 |
| DeepSeek | 25.0 → 35.0 | 35.0 → 38.0 | 8.0 → 7.0 | 3.0 → 5.0 | 25.0 → 10.0 |
| Panel Middle | 36.0 → 35.0 | 35.0 → 36.7 | 7.0 → 6.8 | 4.5 → 5.0 | 7.7 → 9.5 |
Seat numbers are each model's own percentages as submitted on August 26. One first-round board came back summing to 100.2 rather than 100, so every seat is normalized to sum to 100 before the middle is taken, which is why the bottom row is not always the plain midpoint of the column above it, and the medians are taken column by column, so they do not sum to exactly 100. The seats were not re-run for this update; the Claude Fable 5 seat is the model that Friday's table now lists second, behind its own successor.
Records to date for the four externally graded seats on this panel, each graded against real market settlements: Gemini 87% on 7,014 graded calls · GLM 83% on 3,179 graded calls · Kimi 86% on 3,412 graded calls · DeepSeek 83% on 3,260 graded calls. Recomputed daily; the full scoreboard is public.
These are model estimates, not predictions of fact and not financial advice. Kalshi is a CFTC-regulated exchange for event contracts, 18+ only, and availability varies. Every number in this piece gets graded in public once the market settles: the running record lives on the full graded scoreboard.
More live boards from the same panel: the broader best AI at the end of 2026 board, where the market is not restricted to coding · the top-ranked AI model board, where the settlement leaderboard is a different one entirely · the GPT-6 release-date ladder, which Thursday's launch answered · the Anthropic IPO ladder, where the before-November-1 rung is 82¢ bid · and Kalshi's daily prediction-market hub for what the panel priced this morning.
Hottest Prediction Markets Right Now
- 2028 Democratic presidential nominee · $194M traded
- 2027 Pro Football Champion · $72M traded
- 2028 U.S. Presidential Election winner? · $62M traded
- Pro Baseball Champion · $60M traded
- 2028 Republican presidential nominee · $60M traded
- Bitcoin price at the end of 2026 · $33M traded
Every market above links to our full AI model verdict; browse them all on the OddsShopper prediction markets hub, and see how every settled call actually scored on the full graded scoreboard.
The IPO Ladder Moved The Same Day
Friday was a two-board day for Anthropic on Kalshi. The before-October-1 contract on the exchange's ladder for when the IPO is confirmed, which our page on the Anthropic IPO ladder recorded at 44/46 on August 23, was 3 bid, 4 ask on Friday night after more than 45,000 contracts traded in a day. The before-November-1 rung was 82/85 and the before-December-1 rung 86 bid. The exchange also listed three new rungs, for October 10, 17 and 24, and all three are quoted too wide to read.
The reason arrived late Friday. Reuters reported, citing people familiar with the matter, that "Anthropic is expected to begin marketing its initial public offering in mid-October at the earliest" and to "complete the listing days before the U.S. midterm elections in November." Anthropic "had been expected to make its IPO prospectus public as early as next week"; Reuters says that is now "not expected until late September." The same report has the company finalizing a $15 billion revolving credit facility, with Morgan Stanley, Goldman Sachs, JPMorgan and Citi among the banks on a deal some investors have put at $2 trillion. Anthropic declined to comment.
So the same evening the coding board made Anthropic a seven-in-ten favorite on a benchmark, the IPO ladder took confirmation before October 1 from a 44-cent chance to a 3-cent one. Neither market caused the other; both are traders updating on the same company in the same week. The row to argue with is the before-November-1 rung at 82 bid. Marketing in mid-October at the earliest and a listing days before the November 3 midterms puts the completion somewhere around the last days of October or the first days of November, and the rung's line runs straight through that window. An 82-cent price says the calendar lands on the near side of it. Reuters' wording does not say that.
Where The Panel Changed Its Mind
The revision round moved real numbers on this board, and no money changed hands doing it. The clearest example came from the seat that started furthest from everyone else. Where the seats say FIELD below, they mean the last row of the board: nobody with a contract holds the top score, so every one of the nine pays zero.
"The FIELD deserves high probability (25%) because unlisted organizations like Abacus.AI (82.47) are highly competitive, and the board excludes major players like Meta." — DeepSeek, first round
After reading the other seven arguments, it cut that to 10 and rewrote the reasoning around a narrower mechanism, while moving Anthropic up ten points:
"FIELD deserves 10% because Abacus.AI's fine-tune already ranks fifth, demonstrating unlisted organizations can compete for this narrow column." — DeepSeek, second round
Read that first-round quote again with Friday's table in hand. The seat named Meta as the excluded player to worry about, and Meta's Muse Spark 1.3 was on the next table LiveBench posted, in twelfth.
The most interesting revision was an argument nobody else had made. Four of the eight seats treated Anthropic's July 1 safety-classifier redeployment as a live threat to the leader's score. One seat went back and checked the dates:
"A, C, D and H all treated Anthropic's July 1 classifier redeploy as a live downside on the 85.99. But the timeline in the card forecloses that: Fable 5 was suspended June 12, the question set is dated June 25, and site data was refreshed August 25, so 85.99 is almost certainly the post-classifier score already." — Claude Opus 5, second round
[Editor's note: LiveBench does not publish a run date beside any individual model score, so this could not be confirmed from the settlement source in August. Friday's table settles the practical question a different way: the successor model scored 86.38, above the 85.99 the seats were arguing about.]
And one seat moved with the panel, from 5 to 9 on the unlisted field, but for a reason nobody else gave. Its argument turns on a loophole in who gets credit: anyone can take a strong open-weight model, tune it further, and be listed as the owner of the result.
"I hold FIELD at 9, above most of the panel. Abacus.AI already sits fifth at 82.47 by fine-tuning Moonshot's open weights, LiveBench attributes fine-tunes to the tuner, and any December open-weight frontier drop invites the same arbitrage." — Claude Fable 5, second round
What The Panel Said Would Move It, And What Did
Four things, each of them checkable rather than atmospheric. One of them has now happened, and the one the panel rated most likely has not.
A Codex-branded release from OpenAI. Seven of eight seats named this as the highest-probability crown flip in the window. It has not happened: GPT-6 Astra is not Codex-branded, and it scored 80.36 on the column. A Codex variant of the 6 generation, scored above 86.38, is now the most likely OpenAI path the table shows. Direction if it lands: pushes OpenAI up and everything else down.
A Gemini Pro release before December. Google's price and its leaderboard position are both explained by a Pro line that, as TechCrunch reported on July 21, had last been updated in February, with Bloomberg reporting internal delays on a 3.5 Pro. Friday's table carries a Gemini 3.8 Flash that August's did not, and still no Pro. Direction if it lands: pushes Google up sharply, Anthropic and OpenAI down.
Another Anthropic frontier model that scores above 85.99. This is the one that fired. Fable 5.1 scored 86.38 on September 4, three days after release, and the market repriced it from the last pre-refresh print of 50 to 66 within 15 minutes. Direction, as written in August: pushes Anthropic toward its own ceiling. The ceiling question is now whether 69 to 73 cents already prices a four-month hold.
An open-weight frontier release in the fourth quarter. Abacus.AI reached fifth by fine-tuning someone else's open weights, and LiveBench credits the tuner. Meta's arrival in twelfth is a second unlisted name near the top. Direction: pushes the unlisted field up and every listed company down.
Settlement timeline
| Date | What happens |
|---|---|
| June 25, 2026 | The question set the contract currently settles against |
| September 1, 2026 | Anthropic releases Claude Fable 5.1 and Mythos 5.1 |
| September 3, 2026 | OpenAI releases GPT-6 Astra |
| September 4, 2026, 4:09 P.m. ET | LiveBench refreshes the table: Fable 5.1 first at 86.38, GPT-6 Astra fourteenth at 80.36 |
| December 31, 2026, 10:00 A.m. ET | The snapshot; last trading time is the same moment |
| No Later Than January 1, 2027 | Settlement, unless the outcome goes to review |
This page is re-scored when the story moves. Market prices on it are stamped September 4, 2026; panel numbers are stamped August 26, 2026.
More on this: Where Does Bitcoin End 2026? Kalshi's Price-Band Odds · Nurmagomedov Vs Song Odds: Not SO Fast On Umar · Kalshi Venezuela Leader Odds: Who Holds Power In 2026? · Why Andy Barr (R) Will Win the Kentucky Senate Race · Why Jeff Merkley (D) Will Win the Oregon Senate Race
The Bottom Line
In August the market and the panel were looking at the same leaderboard and reading two different things into it: traders priced the incumbent, the panel priced the pattern of a lead that had changed hands five times. Friday's refresh handed the traders a second incumbent. Anthropic now holds first and second on the column that settles this, its lead over the field grew rather than shrank, and the model OpenAI shipped to answer it landed fourteenth on the only two tasks the contract counts. A quote of 69 to 73 cents says the market has read all of that correctly.
What the quote does not say is how much conviction is behind it. The repricing was done on about 1,250 Anthropic contracts against nearly 130,000 already open, the 73-cent offer is 15 contracts deep, and the real offer for size is 75. This market repriced within 25 minutes of a leaderboard refresh because that is what it trades, and it can move back the same way the next time LiveBench posts a table with a Codex in it. Between now and December 31 there will be more refreshes than headlines, and the refreshes are the only ones that pay.



