The Best AI Model in 2026: An Expert Analysis of What a $4.2M Prediction Market Reveals
Anthropic leads a $4.2M Polymarket bet on the best AI model in 2026. Analyze what traders' implied odds reveal for developers, founders, and SaaS buyers. Discover.

If you are choosing an AI platform in late 2026, the real question is not simply which company has the best model today. It is whether Anthropic’s apparent benchmark lead is durable enough to justify standardizing on Claude—or whether the rapid turnover at the frontier makes a multi-model strategy safer.
The September Polymarket gives an unusually emphatic answer to the narrower question. As of September 30, 2026, traders price Anthropic at a rounded 100% implied probability of resolving as the company with the best model, while OpenAI, Meta, Google, SpaceXAI and DeepSeek each sit at a rounded 0%. Roughly $4.24 million has traded in the market.[1] That is a strong signal about the expected resolution of one leaderboard snapshot, not proof that Anthropic will dominate every workload—or even remain ahead through year-end.
Bottom line
>
- The market implies near-certainty that Anthropic wins the specified September 2026 resolution.
- Traders appear to be rewarding Claude’s combination of benchmark leadership, agentic performance and falling API prices.
- The odds do not mean OpenAI, Google or open-weight models are irrelevant; the contract measures a particular public ranking at a particular time.
- Developers and SaaS companies should treat the market as a sentiment indicator, then validate models against their own tasks, costs and security requirements.
What is the $4.2 million prediction market actually pricing?
The Polymarket contract asks which company will have the best AI model at the end of September 2026. Its resolution rules—and the leaderboard snapshot they depend on—matter more than the broad wording. Traders are not evaluating every possible definition of “best.” They are betting on which company satisfies the contract’s designated ranking mechanism.[1]
At the September 30 snapshot, the reported prices and trading volumes are:
| Company | Implied probability | Reported amount traded |
|---|---|---|
| Anthropic | **100%** | **$1,028,979** |
| OpenAI | **0%** | **$815,651** |
| Meta | **0%** | **$489,547** |
| **0%** | **$434,135** | |
| SpaceXAI | **0%** | **$357,982** |
| DeepSeek | **0%** | **$187,177** |
The total market volume is approximately $4,236,132, including trading beyond the six positions listed above.[1] Third-party market pages also track the contract and its pricing, but Polymarket remains the primary source for the resolution terms and live market data.[3][4]
A displayed 100% or 0% should be read as rounded market pricing, not metaphysical certainty. Prices can compress toward the endpoints when traders believe the relevant evidence is already visible, liquidity is uneven or resolution is imminent.
The volume distribution is informative. More than $1 million traded on Anthropic, but OpenAI attracted more than $815,000 and Meta nearly $490,000. That suggests the market was not always this one-sided: substantial capital accumulated around competing hypotheses before pricing converged.
Prediction markets have increasingly become part of the way AI observers narrate the race:
Over the last few days, on Polymarket, $GOOGL's lead as the best AI model by the end of 2025 has increased even further.
Anthropic odds have also risen, while those of OpenAI and xAI have decreased.
While $GOOGL's market share in LLM search will be much lower than the market share it has on traditional search on the enterprise side if $GOOGL turns out to be the best model provider and on top of it offers them via GCP on their TPU infrastructure, GCP's value could be much more than the market current anticipates.
But odds are sensitive to the exact question. A wager on a September leaderboard is different from a bet on enterprise adoption, revenue, developer mindshare or the strongest model at the end of 2026. Earlier market commentary has similarly turned model odds into broader investment theses:
Jason's AI Pair Trade: Short OpenAI. Long Google, xAI, and Anthropic.
Why? OpenAI's competition is fierce.
"They're facing a Google firing on all cylinders, Anthropic, and Grok beating them in the leaderboards pretty consistently."
Polymarket has Google's Gemini 3 at ~87% to finish 2025 as the top-ranked LLM.
Over the last six months, Gemini has started to shrink ChatGPT's massive lead in traffic share.
That extrapolation can be useful, but it can also go too far. Winning one resolution does not automatically imply winning cloud infrastructure, consumer traffic or SaaS distribution.
Why do traders currently price Anthropic so heavily?
The strongest explanation is release cadence. Anthropic appears to have assembled a sequence of benchmark-leading models rather than relying on one isolated launch.
Artificial Analysis reported that Claude Fable 5.1 led its Intelligence Index before the latest Opus release:
Claude Fable 5.1 tops the Artificial Analysis Intelligence Index but costs 20% more per task than Fable 5 despite a 75% cache read price cut
We supported @AnthropicAI with pre-release evaluation of Claude Fable 5.1. At max effort it scores 66 on the Artificial Analysis Intelligence Index, the highest score we have measured, ahead of Claude Opus 5 (max, 63), Claude Fable 5 (max, 62), GPT-5.6 Sol (max, 61) and Grok 4.6 (high, 61). We evaluated the model with Anthropic's ‘default’ server-side fallback, which routes safety-flagged requests to Claude Opus 4.8 or Claude Opus 5; fallback served ~4% of output tokens across the Intelligence Index.
Key takeaways
➤ Frontier Intelligence with improvements across benchmarks: Fable 5.1 gains +4 points on the Intelligence Index over Fable 5. On HLE, Fable 5.1 scores 59.1%, ahead of the previous best of 55.5% from Claude Fable 5. It posts the narrowly highest scores we’ve seen on Terminal-Bench v2.1 (91.4%) and SciCode (62.0%), and on τ³-Banking it gains 9 points over Fable 5
➤ 75% cache read price cut, but Fable 5.1 still costs more per task: Anthropic has cut the cache read price from $1 to $0.25 per 1M cached input tokens, with standard pricing unchanged at $10/$50 per 1M input/output tokens. Fable 5.1 (max) costs $3.76 per Intelligence Index task, 20% more than Fable 5 (max), because it uses ~1.7x the output tokens. The cache cut saves ~$1.40 per task, concentrated in the agentic evaluations where the majority of input tokens are cache reads. At xhigh effort Fable 5.1 scores 65 at $2.72 per task, $1.04 less than max, but still above Claude Opus 5 (max, 63) at $2.34
➤ Claude Fable 5.1 holds the upper end of the Intelligence vs Output Tokens per Task Pareto frontier: every model variant scoring higher than GPT-5.6 Sol (medium) on the Intelligence Index is matched or beaten by a Fable 5.1 effort level on both intelligence and token usage
➤ Highest scores on agentic work tasks, but effectively tied with Opus 5: Fable 5.1 sets the highest scores we have measured on GDPval-AA v2 (1,853 Elo, +130 over Fable 5) and AA-Briefcase (1,694 Elo, +122 over Fable 5), our agentic knowledge work evaluations. Against Claude Opus 5 the GDPval-AA v2 lead is within the confidence interval and AA-Briefcase (1,685) is effectively tied, with Fable 5.1 ahead on analytical quality and rubric correctness, but behind on presentation
That evaluation also exposed an important qualification: Fable 5.1’s top score came with greater output-token use and a higher cost per task than Fable 5. Benchmark leadership was not synonymous with efficiency.
Claude Opus 5.5, announced September 22, then reportedly moved into the top position while receiving a 20% API price cut relative to Opus 5.[9] Anthropic’s product materials position Opus as its highest-capability model, while external reporting emphasized its gains on agentic and professional-work benchmarks.[7][12]
Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index, along with a 20% price cut and larger cache hit discount
Claude Opus 5.5 brings Anthropic to parity with GPT-6 Astra on evaluations like Terminal-Bench 4.0 and AutomationBench-AA, while extending Anthropic’s lead in agentic knowledge work.
At max effort it scores 58 on the Artificial Analysis Intelligence Index, the highest score we have measured by several points. Anthropic has cut Opus pricing to $4/$20 per 1M input/output tokens (Opus 5: $5/$25) and cache reads from $0.50 to $0.20.
Key takeaways:
➤ Consistent strong performance, with leading scores on six of the ten Intelligence Index evaluations: Humanity's Last Exam 61.4% (previous best 59.1%, Claude Fable 5.1), SciCode 66.9% (63.1%, Fable 5.1), GDPval-AA v2.1, AA-Briefcase v1.1, AA-Omniscience and AutomationBench-AA. On Terminal-Bench 4.0 it scores 59.6%, level with the leader GPT-6 Astra (xhigh) and +11 points over Opus 5. It remains slightly behind on CritPt, AA-LCR, and GDP.pdf
➤ Leads in agentic knowledge work: On AA-Briefcase, our private frontier knowledge work evaluation, it reaches an Elo of 1822. This is +143 over Fable 5.1, ahead on both analytical quality and presentation, and is the first time Anthropic has reached presentation quality surpassing GPT-5.6 Sol. This evaluation tests whether models can produce accurate and well-presented professional outputs using our open source reference agent harness, Stirrup
➤ Level with Opus 5 on cost per task despite 1.6x the output tokens: Opus 5.5 (max) uses ~119k output tokens per Intelligence Index task, against ~73k for Opus 5 (max), ~78k for Fable 5.1 (max) and ~27k for GPT-6 Astra (max)
➤ Four of five effort levels sit on the Intelligence vs Cost per Task frontier: Opus 5.5 max, xhigh, high, and medium all sit on the Pareto frontier, costing less or outperforming other models scoring 50+ (GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5)
For buyers, the pricing change may be as consequential as the ranking. A better model that costs more can remain confined to high-value tasks. A stronger model whose unit price falls can replace cheaper models across coding agents, customer-support automation and document workflows.
Claude Sonnet 5.5 intensified that pressure after its September 28 release. Reporting indicated that it nearly matched Opus 5.5 on benchmarks while costing as much as 30% less per task.[8][10] That gives Anthropic a two-tier proposition: Opus for maximum capability and Sonnet for workloads where latency and cost matter more.
Practitioners on X interpreted the sequence as Anthropic retaking the overall lead after a temporary OpenAI advantage:
Since mid 2025, OpenAI has been cheaper but worse than Anthropic. Recently OpenAI became better, but then with Opus 5.5, Anthropic is now the best overall.
ChatGPT has access to Reddit and Gemini has access to YouTube - neither of which Claude has. Also ChatGPT can do images which Claude doesn’t have. Otherwise Claude is better but more expensive.
With OpenAI DevDay next, I hope the answer is not just another product layer. but a genuinely more intelligent frontier model at a dramatically lower price.
Anthropic has been the most consistent winner in the recent frontier-model race. OpenAI briefly reclaimed the lead during a period when Anthropic’s Methos model was unable to be released, leaving a temporary gap at the top. That advantage did not last: Claude Fable 5.1 essentially matched OpenAI’s frontier score, and Opus 5.5 and Sonnet 5.5 then moved decisively ahead on Artificial Analysis.
The more important signal is economics. Anthropic is now delivering higher measured frontier intelligence at materially lower API prices. OpenAI is being squeezed from both directions. Anthropic above it on capability, while lower-cost challengers such as Muse are closing in from below.
On this benchmark, the trend is difficult to ignore: Anthropic has retaken the frontier crown, while OpenAI is carrying premium pricing without the corresponding intelligence lead.
The market therefore appears to be extrapolating three related signals:
- Benchmark momentum: Fable, Opus and Sonnet releases repeatedly placed Anthropic at or near the top.
- Agentic strength: Claude’s advantage appears concentrated in multi-step coding and knowledge-work tasks that map more directly to paid SaaS use.
- Improving economics: Price cuts reduce the penalty for choosing the frontier model.
That combination matters more than a single score. A model provider becomes strategically dangerous when it can improve capability and cut prices simultaneously.
Are traders discounting OpenAI and Google too aggressively?
A rounded 0% price does not mean traders believe OpenAI or Google has no competitive model. It means the market sees almost no remaining path to those companies winning this specific September contract.
GPT-6 Astra illustrates why that distinction matters. One widely discussed Code Arena result placed Astra Max at 1,797 on the WebDev leaderboard—35 points above Claude Fable 5.1 and more than 100 above Claude Opus 5—at roughly comparable blended pricing:
The benchmark throne just changed hands.
GPT-6 Astra (Max) scored 1,797 on Code Arena's WebDev leaderboard this week, clearing Claude Fable 5.1 by 35 points and Claude Opus 5 by over 100. That's a +180pt improvement over where OpenAI's prior frontier model, GPT-5.6 Sol, sat at #13.
What makes this more interesting than a score flip: the model costs the same as Anthropic's top tier. Both are at roughly $40/Mtoken blended. The Pareto frontier moved without a pricing gap opening up.
My read on the category breakdown: Astra leads in Data & Analytics, Consumer Product, and Content Creation Tools. Claude Fable 5.1 still holds in other categories. The frontier is not monolithic right now.
That result suggests the frontier is workload-dependent. An OpenAI model can lead web development while Anthropic leads a broader composite or agentic-work index. “Best model” compresses multiple dimensions—coding, research, tool use, writing, visual generation, cost and latency—into a label that no single benchmark fully supports.
Google has a different strategic advantage. Gemini can be distributed through Google Cloud, existing enterprise relationships and Google’s TPU infrastructure. The company also controls products and data surfaces that can complement model quality. AI model rankings and real-usage trackers show why capability and deployment reach should be evaluated separately.[13][14]
The practical pattern is closer to a portfolio than a winner-take-all hierarchy:
I wouldn't say that - just now for what I need Claude wins. However:
I use @bot to manage personal stuff its more fun to talk too and handles things good enough
@GeminiApp is better at researching - not at turning researching into result but they have so much data they are good for fact checking things and research when writing
GPT Image 2.5 is best image model
Grok 4.7 is best at solving niche issues in fintech world from my experience and understanding compliance and bank documentation
@ElevenLabs wins in voice
So I am not biased to say @claudeai is best and that can't change. However I would argue currently right now in what I need of AI, which is mostly coding and marketing, Claude wins from my personal experience. Sorry if that made you think "nothing beats Claude"
For a SaaS team, that division of labor can be rational:
- Anthropic may fit coding, complex writing and agentic knowledge work.
- OpenAI may fit image generation or tasks where its broader product ecosystem matters.
- Google may fit research-heavy workflows, long-context data work or organizations already standardized on GCP.
- Specialized providers may outperform general models in voice, finance or other verticals.
The September odds are therefore strongest as a forecast of contract resolution and weaker as a forecast of customer behavior.
Does the public leaderboard reflect the real AI frontier?
Prediction markets can price only observable events. AI labs, however, may possess internal systems that are ahead of what they have released.
One X observer summarized the core uncertainty:
I really need you to internalize this:
- the current public frontier is in terms of historical progress 3-6 months behind the private frontier
- most benchmarks are still single-agent and only using a few million tokens, while the latest frontier models are trained for multi-agent operations
OpenAI and Anthropic are both 1.5-2 model iterations ahead, meaning something like GPT-6.1-Astra and Mythos 5.2
they are continuing to race internally
The specific three-to-six-month estimate is not independently established by the supplied reporting, but the conceptual distinction is important. A lab may delay a model because it lacks inference capacity, considers deployment too expensive, has not completed safety testing or cannot serve it reliably at customer scale.
Discussion around Anthropic’s rumored Mythos system shows how compute enters the release decision:
anthropic doesn't have enough compute to publicly release mythos
the api pricing also suggests it could be far larger than gpt-5.5 base model
anthropic has always reached the frontier by using bigger models and more tokens -- while openai focuses more on efficiency and serving billions of users
This produces two races:
- The private frontier concerns the strongest system a lab can run internally.
- The commercial frontier concerns the strongest system it can expose through a stable, affordable API.
Polymarket’s September contract prices the second race. If Anthropic has a stronger internal model but cannot release it, that model should not determine the public resolution. Conversely, a theoretically weaker model can win commercially by being available, affordable and reproducible.
For SaaS buyers, shipped capability is the correct priority. An unreleased model cannot meet a service-level agreement, pass a procurement review or support a production agent. But founders should still monitor the private-frontier discussion because delayed systems can abruptly change pricing and product roadmaps when they arrive.
Can buyers trust the benchmarks behind “best model” markets?
Benchmark methodology is the largest source of fragility in this kind of contract.
Artificial Analysis revised its Intelligence Index to version 4.2, changing benchmark composition, held-out weighting and grading infrastructure. One observer noted that the revision separated GPT-6 Astra from GPT-5.6 Sol after the earlier index had tied them:
I'm glad Artificial Analysis is updating its Intelligence Index
In the previous version, GPT-6 Astra and GPT-5.6 Sol both scored 61. That was at least suspicious. Astra is a newer frontier model, so seeing it tie OpenAI's previous flagship made the leaderboard look wrong
In v4.2, Claude Fable 5.1 leads at 57, Astra is second at 55, and Sol sits at 51. This looks much more believable
AA added private AA-Briefcase, Surge's GDP.pdf, removed saturated GPQA Diamond, doubled held-out weighting to 40%, and improved parts of the grading infrastructure
Astra's profile makes more sense now too: ~85 Elo above Sol on AA-Briefcase and 33.2% vs 28.2% on GDP.pdf. Honestly, this is the part I like most: AA revisiting the index instead of treating the old ranking as final
Updating a benchmark is healthy; treating an old ranking as immutable would be worse. But revisions demonstrate that a leaderboard is a measurement system, not ground truth. Change the tests or weights and the winner can change.
Three methodological issues deserve attention:
- Composite weights encode priorities. A leaderboard may emphasize scientific reasoning, terminal use or professional documents differently from an actual SaaS workload.
- Token and agent budgets affect results. A model allowed to produce substantially more tokens—or operate through a richer agent harness—may score higher while costing more and responding more slowly.
- Scores contain uncertainty. Small leads may fall inside confidence intervals, especially on subjective or model-graded tasks.
The gap becomes wider as frontier systems move toward multi-agent operation. Most public evaluations remain easier to reproduce when they use bounded token budgets and standardized single-agent harnesses. Those constraints may under-measure systems designed to delegate, retry and coordinate tools.
The right procurement response is not to ignore public benchmarks. It is to use them for shortlisting, then run a private evaluation that measures:
- success rate on representative tasks;
- total cost per successful outcome;
- latency at peak traffic;
- tool-call and structured-output reliability;
- security and data-retention requirements;
- frequency and severity of unacceptable failures.
A two-point public lead is irrelevant if the runner-up is cheaper and more reliable on your application.
How do GLM 5.3 and DeepSeek change the cost equation?
The September market gives DeepSeek a rounded 0% probability, but that says little about the structural impact of open-weight models. Open weights can matter without taking first place on a designated frontier leaderboard.
Merge reported that GLM 5.3 beat Claude Sonnet 5 across its set of 20 coding tasks at one-tenth of Claude’s cost:
We tested 5 open-weight models against @AnthropicAI Claude Sonnet 5 on 20 real coding tasks.
• @Zai_org GLM 5.3
• @Zai_org GLM 5.3 Flash
• @deepseek_ai DeepSeek V4 Pro
• @deepseek_ai DeepSeek V4 Flash
• @Kimi_Moonshot Kimi K3
GLM 5.3 won at a 1/10 of Claude's cost:
That is one organization’s evaluation, not a universal ranking. Still, it reframes the buying question from “What is best?” to “What is good enough at the lowest total cost?”
For a startup making thousands of daily calls, a modest quality gap may be acceptable if an open model cuts inference expense, can be fine-tuned and keeps sensitive data inside controlled infrastructure. Open weights also provide leverage during contract negotiations with frontier vendors.
Control is increasingly part of the argument:
If you’re a researcher who uses Claude Code or Codex for your daily work, consider using open models instead.
Recent events have shown why owning the entire stack is so important. While OpenAI and Anthropic currently offer the strongest models, using them means working on their terms. OpenAI and Anthropic have shown they are not afraid to alter model output, service, and access if user interests conflict with their business interests.
OpenAI:
- Sep 2026: accused of using Codex data from Buckmaster and Alpöge to race to a solution to Navier-Stokes using their massive compute advantage. OpenAI later admitted that they “cannot rule out that de-identified data derived from their usage of our products helped improve our models”.
- Aug 2026: announced removal of all OpenAI models on Cursor after their SpaceX acquisition
- Dec 2025: injected ads into ChatGPT conversations, even for users paying $200/mo subscriptions
Anthropic:
- Jun 2026: launched Fable 5 with safeguards that limit Claude’s effectiveness at ML research tasks through interventions that are not visible to the user
- Apr 2026: removed subscription coverage for third-party tooling such as OpenClaw and Pi
- Jan 2026: cut off xAI engineers’ Claude access in Cursor
The only way to protect yourself from these kinds of interventions by the labs is to own the model, the tooling, and the data. This is particularly important for researchers, who often work on confidential projects and with sensitive data.
You should do research on your terms, with tools you control and work that remains yours.
Owning the stack does not eliminate risk. It transfers responsibility for deployment, patching, observability, model evaluation and abuse prevention to the operator. That makes self-hosting a better fit for teams with:
- meaningful and predictable inference volume;
- experienced ML infrastructure staff;
- stringent privacy or sovereignty requirements;
- a need to modify weights or inference behavior;
- enough scale to amortize GPUs and operational complexity.
Small teams with variable traffic are generally better served by managed APIs and a provider-neutral integration layer.
Open weights also introduce security externalities. One X post discussing an Anthropic red-team assessment alleged that GLM-5.3 approached a held-back Anthropic model on exploit construction:
Frontier-level cyberattack capability just went open-weight and downloadable.
Anthropic's red team tested Zhipu's GLM-5.3 and found it nearly matches Claude Mythos Preview at building real exploits. On ExploitBench: 50/410 vs. 56/410. Close enough to matter.
The scarier stat: a $20 API run, 20 minutes of human attention, and GLM-5.3-Flash turned a freshly disclosed Chrome bug into a working attack that bypassed processor-level security.
Anthropic held Mythos back specifically to give defenders a head start. GLM-5.3 ships with no comparable guardrails and no access controls.
Fair caveat: Anthropic has obvious reasons to sound this alarm loudly.
Does CAISI's "four months behind the best US models" framing actually mean anything useful for policy, or does it collapse the moment the next open-weight drop lands?
The poster also acknowledges Anthropic’s commercial incentive to emphasize that risk. Even so, the underlying tension is real: the same availability that reduces vendor control can make capable cyber systems downloadable and difficult to gate.
Open-weight adoption will therefore push SaaS security reviews beyond data privacy. Buyers will need to ask who patches the model-serving stack, monitors abuse and bears responsibility when an autonomous system generates harmful actions.
What should developers, founders and SaaS buyers do with these odds?
Prediction markets can aggregate information quickly, but they cannot replace technical evaluation. Experiments in asking AI systems themselves to construct prediction-market portfolios illustrate both the appeal and the circularity:
Best AI models bet $1000 on polymarket
asked to apply modern portfolio theory and bet sizing to create a calculated portfolio
they can jump into literally everything
expected profits:
> ChatGPT 5.2 +47%
> Grok 4.1 +1464.8%
> Gemini 3 pro +21.5%
> Claude +28.6%
wanna show that system is more important than intuition
even if this system consists of ai that sometimes output complete nonsense
i share the prompt and methodology so anyone can repeat it
I'm using the best AI models to bet $1000 on Polymarket!
Asked it to use modern portfolio theory + bet sizing to make calculated bets. It chose everything from BTC price to Fed rates.
Expected returns:
o3-pro: +21.6%
opus 4: +41.7%
grok 4 heavy: +34%
Will report back who won.
The actionable lesson differs by audience.
Developers should avoid hard-coupling products to one model
Use a model gateway or internal abstraction that normalizes authentication, tool calls, retries, structured output and logging. Keep prompts versioned by model because identical prompts rarely transfer perfectly.
For coding agents or complex knowledge work, the market and recent benchmarks support making Anthropic the first model to evaluate. Keep at least one OpenAI or Google fallback, and test an open-weight option if cost or data control matters.
Founders should optimize cost per completed task
Token prices alone are misleading. A cheaper model that needs repeated attempts can cost more than an expensive model that succeeds immediately. Conversely, a benchmark winner that emits 1.6 times as many tokens may damage gross margin.
Anthropic’s reported Opus and Sonnet price moves indicate intensifying competition.[9][10] Buyers should expect providers to use model tiers, caching discounts and effort controls to compete for production volume. That favors architectures capable of routing simple tasks to cheaper models and difficult tasks to frontier systems.
Enterprise SaaS buyers should hedge operationally
Enterprises should negotiate portability, data-export rights and clear policies on model retirement. They should also require regression testing before silent model upgrades.
A practical deployment might use a managed frontier model as the default, a second vendor for continuity and an open-weight model for sensitive or high-volume workloads. This is more complex than choosing one provider, but it reduces exposure to outages, access-policy changes and sudden pricing shifts.
How should the September signal be read without over-reading it?
The strongest conclusion is narrow: the market implies near-total confidence that Anthropic satisfies the September 2026 contract’s definition of the best AI model. The pricing is consistent with Anthropic’s benchmark momentum, agentic-work performance and lower API costs.
The broader race remains less settled. A separate year-end market has priced Anthropic around 78%, showing materially more uncertainty when traders extend the horizon beyond September.[5] More time creates more opportunity for an OpenAI, Google or open-weight release to alter the rankings.
Earlier research comparisons also show why durable conclusions require caution: model advantages can vary sharply by task, and even strong AI systems may remain well below expert humans on difficult research work.
Anthropic Beats OpenAI in AI Research Tests
Via the Information
In a first-of-its-kind evaluation by the nonprofit METR, Anthropic’s advanced AI model, Claude Sonnet 3.5, demonstrated superior performance in conducting AI research compared to OpenAI’s o1-preview. Out of seven challenging tasks, Claude excelled in five, delivering particularly strong results in two. OpenAI’s model won in two other tasks, with one being a decisive victory.
While both models showed impressive capabilities, they fell short when compared to human researchers, who scored more than double the average of the AIs. However, Claude matched human performance on two tasks, and o1-preview achieved this in one. The problems tested required high levels of creativity, hypothesis generation, and experimental design, such as writing a language model without using division or exponents.
The best posture for practitioners is therefore:
- Evaluate Anthropic first for coding and agentic knowledge work when quality is the priority.
- Keep OpenAI and Google available for modality, research, ecosystem or cloud-distribution advantages.
- Maintain an open-weight fallback when privacy, customization, continuity or unit economics justify the operational burden.
- Re-run private evaluations after major releases, rather than treating any September leaderboard as permanent.
- Use prediction markets as a leading sentiment indicator, not as an architecture diagram.
The $4.2 million market is telling the industry that traders believe Anthropic has won this snapshot. The behavior of developers tells a more durable story: the frontier changes too quickly for sophisticated buyers to bet their entire stack on one lab.
Sources
[1] Which company has the best AI model end of September? Trading Odds & Predictions 2026 — Polymarket
[3] Which Company Has the Best AI Model in September 2026? — Lines.com
[4] Which company has the best AI model end of September — CryptoSlate
[5] Which company has best AI model end of 2026? — Polymarket odds
[9] Introducing Claude Opus 5.5 — Anthropic
[10] Claude Sonnet 5.5 nearly matches Opus 5.5 while costing less per task — The Decoder
[12] Anthropic releases Claude Opus 5.5 — VentureBeat
[13] AI Model Rankings — AIgateway
[14] Who Is Winning the AI Race? Monthly LLM Leader Timeline — BenchLM.ai
References (15 sources)
- Which company has the best AI model end of September? Trading Odds & Predictions 2026 | Polymarket - polymarket.com
- Best AI Model of 2026 Odds: Who Will Be #1 at Year-End? - defirate.com
- Which Company Has the Best AI Model in September 2026? Winner Odds | Lines.com - lines.com
- Which company has the best AI model end of September Odds & Prediction Market Analysis | CryptoSlate - cryptoslate.com
- Which company has best AI model end of 2026? · Anthropic 78% (+3pp) — Polymarket odds - pdata.world
- Polymarket: Anthropic tops best AI model odds at 95.5% on $6.26M volume - blockchain.news
- Claude Opus \ Anthropic - anthropic.com
- Claude Sonnet \ Anthropic - anthropic.com
- Introducing Claude Opus 5.5 \ Anthropic - anthropic.com
- Anthropic's Claude Sonnet 5.5 nearly matches Opus 5.5 on benchmarks while costing up to 30 percent less per task - the-decoder.com
- Best Anthropic Models (September 2026) — Ranked by Benchmark Data | BenchLM.ai - benchlm.ai
- Anthropic releases Claude Opus 5.5, beating Fable 5.1 on key agentic benchmarks at 60% cheaper API price - venturebeat.com
- AI Model Rankings — frontier models by benchmark + real usage · AIgateway - aigateway.sh
- Who Is Winning the AI Race? Monthly LLM Leader Timeline (September 2026) | BenchLM.ai - benchlm.ai
- AI model rankings, prices and new releases — Artificials - artificials.net