market-watch

The Best AI Model in 2026: Why Traders Put 98% Odds on Anthropic — An Expert Analysis

Anthropic sits at 98% on Polymarket's best AI model market with $3.9M traded. Discover what these odds reveal for developers, founders, and SaaS buyers. Learn more.

👤 📅 September 23, 2026 ⏱️ 23 min read
AdTools Monster Mascot reviewing products: The Best AI Model in 2026: Why Traders Put 98% Odds on Anthr
How we research: This guide is compiled by the AdTools team from the linked sources below and current public discussion. Pricing and features change often, so please verify time-sensitive details with each vendor before making a decision.

The practical question behind Polymarket’s “best AI model” market is not simply whether Anthropic will lead a leaderboard on September 30, 2026. It is whether developers, founders, and SaaS buyers should treat Anthropic’s apparent lead as a durable platform advantage—or as a short-lived snapshot before the next OpenAI, Google, or open-source release.

The market’s answer is unusually decisive in the short term. As of September 23, traders imply a 98% probability that Anthropic will hold the qualifying top position at the end of September. But longer-dated markets are substantially less certain, suggesting that traders see Anthropic as the clear favorite for this resolution window—not the guaranteed winner of the broader AI race.

Bottom line

>

- Polymarket traders currently price Anthropic at 98%, versus OpenAI at 1% and rounded 0% probabilities for Meta, Google, SpaceXAI, and DeepSeek.[1]

- The likely catalyst is Claude Opus 5.5: a newly shipped model positioned around stronger agentic performance, lower prices, and faster output.[8]

- The market is pricing a specific public leaderboard on a specific date, not long-term market share, private model capability, or the best model for every workload.

- For builders, the stronger signal is that frontier intelligence is becoming cheaper while reliable agent execution—not raw benchmark IQ—is becoming the scarce capability.

What is the September 2026 prediction market actually pricing?

As of September 23, 2026, roughly $3,889,898 has been traded in Polymarket’s market on which company will have the best AI model at the end of September. The market is expected to resolve around September 30 according to its published rules and designated model-ranking mechanism.[1] Secondary market pages also track the same contest and its changing prices.[2][3]

The current implied probabilities and reported company-level trading volumes are:

CompanyImplied probabilityReported traded volume
**Anthropic****98%****$894,819**
**OpenAI****1%****$739,400**
**Meta****0%****$451,830**
**Google****0%****$363,925**
**SpaceXAI****0%****$349,820**
**DeepSeek****0%****$182,177**

These displayed percentages are rounded, which is why they should not be interpreted as mathematically exact or as saying that every company shown at 0% has literally no chance. A prediction-market price is the market’s current collective estimate, shaped by available information, liquidity, positioning, and the contract’s resolution rules. It is not a guarantee.

The most revealing comparison is across time horizons. One market watcher highlighted that Anthropic was priced near certainty for September but materially lower for December:

Prediction Bubbles 🫧 @predictionbubbl Sep 21, 2026

@Polymarket puts @AnthropicAI at 98.6% for best AI model on 30 September. @Kalshi has Claude at 72.5% for 31 December.

Ten days out the market is nearly certain. Three months out it is not.

@OpenAI is 0.4% for September and 13.4% for December.

https://predictionbubbles.net

View on X

That gap contains the central signal. Ten days from resolution, traders mostly need to assess shipped products and known leaderboard results. Three months out, they must price unreleased models, regulatory delays, compute availability, and unpredictable launch timing. A separate end-of-2026 market put Anthropic near 74%, far below the September contract’s 98%.[6]

The September odds therefore imply confidence in a near-term outcome, while longer-dated pricing implies that traders expect the race to reopen.

Why do traders currently put 98% odds on Anthropic?

The immediate catalyst is timing. Anthropic introduced Claude Opus 5.5 on September 22—just over a week before the expected market resolution.[8] Recent reporting described the release as combining lower pricing with performance around Anthropic’s Fable tier, giving traders a shipped and measurable product rather than a roadmap promise.[10]

The X conversation rapidly focused on the model’s benchmark table:

st1ne @SolSt1ne Sep 22, 2026

Anthropic just shipped Claude Opus 5.5 and didn't bother with a press tour - just dropped a benchmark table against everything else on the market

The shift here isn't a new modality or a new trick. It's the same model, still faster, still cheaper, now just further ahead on the numbers that actually predict whether an agent finishes the job

here's the breakdown:

1 - agentic coding, three ways → 66.4% on Terminal-Bench 4.0 (Astra 57.9%, Fable 5.1 55.8%, Sol 37.3%), 54.4% on FrontierCode v1.1, 57.8% on CursorBench 4.0. Not a one-benchmark fluke - it leads on every coding harness that got tested

2 - knowledge work → 1846 on GDPval-AA v2.1, ahead of Fable 5.1 (1735), Opus 5 (1708), and well clear of GPT-6 Astra (1542) and Sol (1588)

3 - reasoning → 67.7% on Humanity's Last Exam with tools, the top score on the table, edging out Fable 5.1 (65.6%) and beating Astra by 10.5 points (57.2%)

4 - computer use → 81.8% on OSWorld 2.0 (partial), ahead of Fable 5.1 at 80.7% - Astra didn't even report a number here

5 - the honest catch → it's not a clean sweep either. Astra actually wins two categories outright: AutomationBench business workflows (41.4% vs Opus 5.5's 40.0%) and Terminal-Bench-Science (64.6% vs 58.7%). Worth knowing before you pick a model by vibes

6 - the stack worth building → pair Opus 5.5 with Jev (TypeSafe AI's new "System One" model, out Sept 15). Jev doesn't write text - it takes unstructured input and returns a typed decision with a calibrated probability in 70-500ms, at a fraction of a cent per call. So Opus 5.5 does the actual thinking - the agentic coding, the long research runs, the judgment calls - and Jev sits inside the loop doing the thousand small classifications and routing decisions the agent doesn't need a full model turn for. System 2 for the hard parts, System 1 for everything that repeats

7 - why this matters → most teams are running one model for both jobs right now - reasoning through a plan AND deciding "is this ticket urgent, yes or no" with the same expensive call. That's the gap Jev is built for, and it's the gap Opus 5.5's own numbers make obvious: the model is good enough that burning it on trivial decisions is waste

Full benchmark table and the Jev docs are worth five minutes before you touch your stack

View on X

The numbers in that post help explain the market position. Opus 5.5 is described as scoring 66.4% on Terminal-Bench 4.0, compared with 57.9% for GPT-6 Astra, while also leading the cited comparisons on FrontierCode, CursorBench, knowledge work, tool-assisted reasoning, and partial OSWorld testing. It was not presented as a clean sweep: Astra reportedly led selected business-workflow and science-terminal tests.

That nuance matters. Traders do not need to believe Anthropic dominates every task. They only need to believe its qualifying model is most likely to top the market’s designated ranking at resolution.

Price and speed strengthened the launch case. Anthropic and recent coverage positioned Opus 5.5 as 40% cheaper than Opus 5 on typical workloads and capable of 30% faster output, with cheaper cache reads.[8][10]

Vaibhav Sisinty @VaibhavSisinty Sep 22, 2026

Anthropic dropped Claude Opus 5.5 and the price-to-performance ratio is the best they've ever shipped.

Fable 5.1 level performance. 40% cheaper than Opus 5. 30% faster output. Cache reads 60% cheaper at $0.20 per million tokens.

At default effort it beats GPT-6 Astra on FrontierCode at 20% of the cost per task.

First model released since Dario called for slowing down. Tested by METR and Frontier Design before launch. Best alignment score of any model they've tested.

What early testers did with it:

→ 680,000-line code migration in less than a day. Weeks of engineering team work.

→ 200,000-line codebase audited and fixed in 3 hours. Opus 5 took 20+ hours for the same job.

→ Translated HAProxy from C to Rust. Passed nearly all regression tests. 51% cheaper than Fable.

Communication got a major upgrade. Writes the way early testers do. Less jargon. Important information first.

One tester preferred its rewrite of their own prompt over the original.

Pricing: $4 input. $20 output per million tokens. Plus increased limits on Pro, Max, and Team plans with a rate limit reset you can save for whenever you need it.

Anthropic said they'd pace the frontier. Opus 5.5 shows what that looks like in practice.

Smarter. Cheaper. Safer.

View on X

The market may also be rewarding the release’s risk profile. Traders can evaluate an available model, published benchmark claims, pricing, and external launch coverage. By contrast, an unreleased rival requires multiple uncertain events: the model must ship before the cutoff, qualify under the rules, become measurable, and outperform the incumbent.

Practitioner reaction reinforced the perception that the improvement was not merely another score increase:

Daniel Morgan @DanPMorgan Sep 23, 2026

I'm sure we'll figure out some quirks with the model, but right now I'm super impressed by the new Claude Opus 5.5.

Previous models often felt like smart knowledge workers, but with little experience. Therefore they have to try a bunch of different approaches just to settle on one that eventually worked. This feels like an experienced and smart worker who instantly knows the right kinds of things to check and which approaches to try first. I think Anthropic have genuinely done an effective lossless intelligence compression here. 🤯

View on X

That “experienced worker” framing points toward what traders may increasingly interpret as model quality: fewer failed approaches, better initial decisions, and higher task-completion rates.

Is Anthropic’s real advantage coding and agentic reliability?

The market may effectively be defining “best” as best at the high-value work enterprises are currently trying to automate. In 2026, that increasingly means coding agents, computer use, tool calls, and long-running workflows rather than isolated question answering.

An agentic model does more than return one response. It plans, invokes tools, observes results, corrects errors, and continues until a goal is reached. Small differences in judgment can compound over dozens of steps. A model that appears only slightly better on a static test may be substantially more useful if it avoids loops and recovers from mistakes.

Ji-Ha @Ji_Ha_Kim Dec 4, 2025

Anthropic's models are really special, like GPT/Gemini/DeepSeek are smarter, better at math, etc. but Claude is the only one who can actually tackle a new coding task it's not familiar with for 20 minutes, iteratively self-correct and finish the task, others get stuck in loops

View on X

That observation captures a distinction conventional benchmarks struggle to measure. A model can be excellent at mathematics or short-form reasoning yet fail when asked to inspect an unfamiliar repository, revise a plan after an error, and preserve context across a 20-minute workflow.

The X discussion attributes Claude’s position partly to data quality, code-focused training, alignment work, and the fit between the model and Claude Code’s surrounding agent scaffold:

Minh Nhat Nguyen @menhguin Aug 28, 2025

Claude models punch above their weight(s) in navigation + Anthropic is v particular abt data quality in all things and is very locked into codegen + good alignment science + Claude Code was built to be used by employees themselves, and the scaffold fits the agent very well.

View on X

The scaffold matters because a production agent is not just a model. It includes prompts, tool definitions, memory, permissions, retry policies, validation, and the execution environment. A model optimized alongside that system can outperform a nominally smarter model placed in a weaker harness.

One widely circulated report even claims that DeepMind engineers use Claude internally and resisted losing access:

tae kim @firstadopter Apr 20, 2026

Anthropic used more coding data in their training runs, so Claude is better at coding. DeepMind now knows this and sees agentic coding taking off exponentially in terms of revenue, so they will use more coding data for future Gemini models to catch up in capability.

"DeepMind engineers use Claude as a daily tool. Most of the rest of Google does not. When the question of equalizing access came up internally, the proposed response was to remove Claude for everyone — which DeepMind objected to so strongly that several engineers reportedly threatened to leave."

View on X

That is not independent proof of overall model supremacy. It is, however, indicative of the practitioner perception traders may be absorbing: Claude’s coding utility is strong enough to be valued inside a major competing AI organization.

For SaaS teams, the implication is straightforward. If the workload is long-horizon software engineering, repository navigation, migration, or debugging, Anthropic currently fits the market’s preferred profile. If the workload is a narrow classification, low-latency completion, or specialized scientific task, the overall leaderboard leader may not be the right operational choice.

Why does OpenAI have only 1% despite remaining a frontier contender?

A naive reading of the market would be that traders think OpenAI is commercially irrelevant. That is not what the contract measures.

The 1% implied probability concerns OpenAI’s chance of winning under one market’s rules at one month-end snapshot. It says little directly about API revenue, consumer distribution, enterprise adoption, research quality, or whether an OpenAI model will be preferable on a specific workload.

The current competitive move may be less a dramatic intelligence leap than a price reset:

TJT @thijos331 Sep 22, 2026

two frontier labs, less than two hours apart, shipped the same product: same intelligence, lower price

Anthropic: Opus 5.5 at 40% less than Opus 5 on typical workloads
OpenAI: Sol and Luna at 50% lower API prices than 5.6, and Artificial Analysis has the index scores level with 5.6, better on some evals, worse on others

that's a price war, not a capability race. good news if you build on top of this, a harder conversation if you are financing the capacity

View on X

OpenAI’s recent lineup has been discussed as a tiered response to Anthropic, including models optimized for heavy work and lower-cost volume use:

Mikadzyki🌙 @Mikadzyki_NFT Jul 1, 2026

OPENAI RELEASED GPT-5.6 AND BEAT ANTHROPIC'S BEST MODELS

OpenAI unveiled a new generation of its models and immediately set a new bar in coding, cybersecurity and biology. The lineup has three models, essentially an answer to Anthropic's Haiku, Sonnet and Opus:

> Sol, the flagship for heavy coding, cybersecurity and biology
> Terra for everyday work with a balance of price and quality
> Luna for fast and cheap high-volume tasks

The flagship has an ultra mode that runs several agents in parallel and splits one complex task between them

On the numbers, Sol leads:

> 91.9% on Terminal-Bench against 88% for Claude Mythos 5
> the only model to cross 50% on Agent's Last Exam
> up to 750 tokens per second at launch in July

But you can't touch it. Access is open to just a couple dozen vetted organizations, while the public release is frozen under a Trump executive order that puts frontier models through a government review before they ship

History repeats itself. Exactly as it happened with Anthropic's Fable 5, the most powerful model of the generation is out, yet it stays out of reach for ordinary users

View on X

The post also highlights an important distinction between existing and accessible. A powerful model restricted to vetted organizations—or delayed by review—may not satisfy the public availability or ranking conditions that matter for the prediction contract. The market can only price a qualifying outcome.

For developers, this is encouraging. When two frontier labs cut prices while maintaining broadly similar intelligence, application-layer companies receive more capability per dollar. For infrastructure investors and companies committed to expensive capacity, it creates margin pressure.

It also changes the metric that matters. Token price is only the beginning; teams need cost per successful task. That includes retries, human review, latency, tool failures, context consumption, and the cost of correcting bad actions.

Shinra @werksiz Sep 21, 2026

Here's the proof:

OpenAI: $20/month ChatGPT Pro, you pay per token on API, forced into their ecosystem.

Anthropic: Fable 5.1 at $0.15/1M tokens. Same quality on most work. Developers actually building production systems picked Fable because the math makes sense.

OpenAI's benchmarks improved 15% last quarter. Fable's adoption grew 60%.

One metric everyone ignores: Cost per successful production deployment.

OpenAI optimizes for "fastest on MMLU." Anthropic optimizes for "developers actually ship with this."

When you're building agents, routing models, agentic systems — you pick the one that doesn't crater your margin. That's Anthropic.

OpenAI's brilliant at research and marketing. Anthropic's brilliant at what actually gets used.

By 2027, enterprise adopts whoever has lower TCO on their actual workload, not whoever has the highest benchmark score.

OpenAI's still playing benchmark chess.

Anthropic's playing business checkers.

View on X

The specific adoption claims in that post should be treated as part of the live practitioner debate rather than as audited market data. But the proposed metric is right. A $1 run that succeeds is cheaper than a $0.40 run requiring three attempts and manual repair.

Do the odds reflect the real AI frontier—or only the public one?

Prediction markets can price only what traders can observe and what the contract can resolve. That creates a structural gap between the public frontier—models people can access and benchmark—and the private frontier inside labs.

Lisan al Gaib @scaling01 Aug 27, 2026

I really need you to internalize this:
- the current public frontier is in terms of historical progress 3-6 months behind the private frontier
- most benchmarks are still single-agent and only using a few million tokens, while the latest frontier models are trained for multi-agent operations

OpenAI and Anthropic are both 1.5-2 model iterations ahead, meaning something like GPT-6.1-Astra and Mythos 5.2

they are continuing to race internally

View on X

Claims about unreleased internal systems are inherently difficult to verify. Still, the underlying logic is sound: model development does not stop on release day. OpenAI and Anthropic are training, post-training, evaluating, and serving successor systems while their current public models are being benchmarked.

Compute constraints can further separate capability from availability:

Haider. @haider1 May 26, 2026

anthropic doesn't have enough compute to publicly release mythos

the api pricing also suggests it could be far larger than gpt-5.5 base model

anthropic has always reached the frontier by using bigger models and more tokens -- while openai focuses more on efficiency and serving billions of users

View on X

If Anthropic has a stronger internal model but insufficient capacity to release it broadly, that model may matter strategically without affecting the September contract. Conversely, OpenAI’s efficiency and distribution could produce greater commercial reach even if Anthropic tops a designated leaderboard.

This helps explain why the market can imply 98% for September while putting materially lower odds on Anthropic at year-end. The short-term contract is a leaderboard bet; the longer-term contract is a release-cadence bet.

The X conversation increasingly describes the public frontier as a two-company race:

🍓🍓🍓 @iruletheworldmo Jun 29, 2026

i get why people want to root for “open source”.

but the distance between openai/anthropic and anything else is gargantuan. and it isn’t only open source that’s miles back, the other closed for-profits are too.

google, meta and xai are nowhere near. only two labs are sitting at the actual frontier, and the government keeps telling you which two: it force-pulled anthropic’s two best models overnight, and made openai submit its newest one to user screening before it would let it ship. it’s doing that to no one else, because there’s nothing else worth controlling.

and even if we only look at the publicly available models from these two, they dwarf anything held back privately by any company on the planet.

whilst mythos feels like another paradigm shift, it’s the result of pushing the scaling laws further than anyone else can. people misunderstand scaling as one single axis to push, when there’s so much left to scale across all of them: pre-training compute, post-training and rl, test-time compute, data.

you’ll start seeing mythos like jumps every two months, opus 4.7 to 4.8 was already about that and 5.5 to 5.6 runs on the same clock, as we’re now deep inside a hard, fast, and turbulent take off scenario.

so as all the best models say, buckle up buttercup.

View on X

The market broadly reflects that view, but it should not be treated as permanent. A company priced at a rounded 0% for September could become competitive after a single release. One forecasting account, for example, priced Google’s October prospects above the contemporary Polymarket price:

PrecisionAlgorithms @precisionalgo Sep 18, 2026

Google the best AI model at end of October? Venue yes at 7.5. We read 18.

Will Google have the best AI model at the end of October 2026?

Polymarket: 7.5% YES
Precision: 18% YES
Gap: +10.5 points

YES = Google is on top at month end.

View on X

That disagreement is useful. Prediction markets aggregate expectations, but they do not eliminate model risk, thin-market distortions, or differentiated research.

Why are open-source models priced at 0% even as they become cheaper and better?

DeepSeek’s rounded 0% September probability does not imply that open models have no commercial value. It implies that traders currently see very little chance of DeepSeek winning this specific frontier-model resolution.

For many startups, frontier leadership is less important than deployability and unit economics:

Deedy @deedydas Sep 10, 2026

DeepSeek might be the new king of open-source models. It obliterates GLM 5.3 and Kimi K3 on benchmarks while being 4x and 10x cheaper. An insane $0.3/M in, $1.2/M out.

And it’s #6 amongst humans on Codeforces!

This is exactly what startups with coding products needed.

View on X

An inexpensive open or open-weight model can be the better choice when a company needs data control, customization, on-premises deployment, predictable high-volume costs, or freedom from provider rate limits. Those advantages can matter more than winning the hardest general-purpose evaluation.

But the frontier gap remains visible on some emerging tests:

Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞) @teortaxesTex Sep 12, 2026

Unhinged, brutal chart for open models.
GPT 5.6 Luna: 4/70
DeepSeek V4.1, GLM-5.3, Grok 4.6: tied at 3/70
Kimi K3: 2/70
Qwen 3.8 Max and Qwen 3.8 37B: 1/70
GLM 5.3-Flash: 0731, 0813: 0/70
OAI has hill-climbed this internally already. Anthropic is midway. Google, Meta – starting.

View on X

At the same time, benchmark quality itself is deteriorating. Once common test questions enter training data—or labs optimize specifically against them—the score stops being a reliable measure of generalization. Artificial Analysis reportedly responded by placing 40% of a new index behind private test sets:

Alok @analogalok Sep 5, 2026

Artificial Analysis had to lock 40% of their new index behind private test sets just to get honest signal on frontier models.

4 quick takeaways:

- Qwen 3.8 27b on par with GPT 5.6 Luna and Deepseek v4 pro. Beats the new K2 horizon 375b a23b

- Claude Fable 5.1 holds #1, but GPT-6 Astra is a token efficiency monster

- Long context reasoning is moving to messy real world slop (4,500+ pages of footnotes, charts, tables)

- Google sitting behind Meta muse spark 1.3, SpaceXAI grok 4.3, Moonshot kimi k3 and ziphu glm 5.3 on the leaderboard. genuinely what is deepmind’s play here?

- RIP GPQA Diamond. Labs finally contaminated and gamed it into irrelevance.

The era of testing models on cute multiple choice science questions is over. It’s agentic enterprise grunts or bust now.

View on X

This is why “best AI model” is becoming harder to define. Public multiple-choice benchmarks are giving way to private evaluations, messy documents, long contexts, tool use, and simulations that resemble enterprise work.

Even realistic agent tests do not produce one universal winner:

DAIR.AI @dair_ai Sep 1, 2026

Banger paper from the Qwen team.

If you evaluate agents on anything longer than a single session, this one is worth your time.

(bookmark it)

E-Commerce Bench runs an agent through a simulated 365-day year operating several online stores at once.

18 frontier models are scored across seven dimensions and no single model dominates.

GPT-5.6 Sol earns the most, growing a 100,000 opening stake into 1,431,425, then ranks 16th of 18 on fraud avoidance and trails Fable 5 on operational efficiency.

Among open-weight models, Qwen3.8-Max-Preview leads at 416,252, 38% above GLM 5.2 (high), and shows the strongest learning over the horizon by progressively bargaining suppliers down across repeated orders.

Paper: https://t.co/pX3u6GABA6

View on X

In that reported simulation, GPT-5.6 Sol earned the most money but ranked near the bottom on fraud avoidance, while Anthropic led operational efficiency and Qwen led the open-weight group. A SaaS buyer in fraud-sensitive commerce would rationally weight those results differently from one optimizing revenue growth.

The decision rule is therefore:

What does Anthropic’s 98% probability signal for AI and SaaS?

The market implies three broader industry shifts.

First, the public frontier is consolidating around a small number of labs while capability below that frontier commoditizes rapidly. Open-source teams may remain behind on the hardest agentic tests while still becoming good enough—and cheap enough—for most ordinary SaaS features.

Second, AI pricing is moving downward faster than many application companies’ differentiation is improving. SaaS vendors that merely wrap a model and charge a large per-seat premium will face pressure as equivalent inference becomes cheaper. Durable margins will come from proprietary workflows, customer data, distribution, integrations, evaluation systems, and accountability—not raw access to a model.

Third, the competitive axis is moving beyond parameter counts:

DataDigest @Num_Feed_ Sep 21, 2026

The real divergence between Anthropic and OpenAI is no longer about raw parameter counts: it is about RL architecture and test-time compute search.

OpenAI focused heavily on post-training reinforcement learning to produce long-chain reasoning tokens, while Anthropic optimized prompt fidelity, automated tool evaluation, and enterprise security guardrails.

Both will have gigawatt compute clusters by 2027. The winner will not be who trains the biggest foundation model, but whose reasoning engine has the lowest hallucination rate in enterprise codebases.

View on X

For enterprise deployment, the winning reasoning engine may be the one that follows instructions, uses tools safely, avoids hallucinated changes, and completes work with minimal supervision. Guardrails and task reliability become product capabilities, not compliance accessories.

Prediction markets are consequently becoming a live sentiment instrument for AI. They compress releases, benchmark results, access restrictions, rumors, and trader positioning into a continuously changing probability. But they remain most useful as a measure of expectation, not truth.

Which model should developers, founders, and SaaS buyers choose now?

Developers building coding agents

Anthropic is the market-favored default for long-horizon coding as of September 23, 2026. The public evidence and practitioner conversation emphasize navigation, self-correction, and task completion. But developers should keep prompts, tool schemas, memory, and evaluation suites as provider-neutral as possible.

Use a routing layer so lower-cost models can handle extraction, classification, and repetitive decisions while a frontier model handles difficult planning and code changes.

Founders optimizing runway

Do not interpret 98% as a reason to lock the entire product into one provider. The simultaneous price cuts suggest continued competition. Negotiate flexibility, avoid excessive single-vendor capacity commitments, and measure complete workflow costs rather than token prices.

Open models may fit high-volume or privacy-sensitive features even if betting markets give them almost no chance of winning the frontier title.

SaaS and enterprise buyers

Run a private evaluation based on actual failure modes:

  1. Task-completion rate without human rescue
  2. Cost per successful workflow
  3. Hallucination and unauthorized-action rate
  4. Recovery from tool or API failures
  5. Latency under realistic concurrency
  6. Performance on proprietary documents and code
  7. Auditability, security, and data-retention terms

Public leaderboards should determine which models enter the test—not which model wins the contract.

The correct interpretation of Anthropic’s 98% implied probability is therefore narrow but valuable: traders strongly expect Anthropic to lead the specified public comparison at the end of September 2026. They do not imply that Anthropic will dominate every workload, win the year, or maintain the lead indefinitely.

For practitioners, the actionable signal is to use Claude where its agentic reliability justifies the cost, preserve the ability to switch providers, and reassess after every major release. In this market, a near-certainty ten days out can become an open race three months later.

Sources

[1] Polymarket — Which company has the best AI model end of September?

[2] Polymtrade — Which company has the best AI model end of September?

[3] Lines.com — Which Company Has the Best AI Model in September 2026?

[4] Polymarket — AI Predictions & Real-Time Odds

[5] CryptoSlate — Which company has the best AI model end of September odds and analysis

[6] PData — Which company has best AI model end of 2026?

[7] TipRanks — Odds On: Which company will be able to claim best AI model in September?

[8] Anthropic — Introducing Claude Opus 5.5

[9] Anthropic — Claude Opus

[10] TechCrunch — Anthropic releases Opus 5.5 with lower prices and Fable-level performance

[11] Mashable — Anthropic launches Claude Opus 5.5: benchmarks, pricing and safety