market-watch

Traders Are Betting on Anthropic: What Polymarket's 99% AI Model Odds Reveal for Developers in 2026

Polymarket odds put Anthropic at 99% for best AI model by end of September 2026. Discover what the $4M market reveals about AI and SaaS for developers.

👤 📅 September 25, 2026 ⏱️ 20 min read
AdTools Monster Mascot reviewing products: Traders Are Betting on Anthropic: What Polymarket's 99% AI M
How we research: This guide is compiled by the AdTools team from the linked sources below and current public discussion. Pricing and features change often, so please verify time-sensitive details with each vendor before making a decision.

The practical question behind Polymarket’s September AI-model bet is not simply whether Anthropic will “win.” It is whether developers, founders, and software buyers should treat a 99% market price as evidence that Claude is now the safest platform choice.

The short answer: no—but the odds are still a meaningful signal. As of September 25, 2026, the market implies that Anthropic is overwhelmingly likely to lead the specific leaderboard used for settlement around September 30. It does not imply a 99% chance that Claude is the best model for every workload, the lowest-cost model to operate, or the long-term winner in enterprise AI.[1]

Bottom line

- Polymarket currently prices Anthropic at 99%, reflecting near-term confidence in a particular leaderboard outcome.

- The same confidence does not extend to December, suggesting traders expect frontier leadership to remain highly flippable.

- Developers should benchmark models against their own tasks; founders should preserve vendor optionality; SaaS buyers should model overages and cost-to-serve.

- The deeper industry shift is from selling a “best model” to controlling the agent, tool, data, security, and deployment layer around it.

The $4 million bet: What does Polymarket actually price on September 25, 2026?

Approximately $4,008,767 has been traded in Polymarket’s “Which company has the best AI model end of September?” market, according to the September 25 snapshot supplied for the market resolving around September 30.[1] The current prices are:

CompanyImplied probabilityVolume traded
Anthropic**99%****$927,812**
OpenAI**0%****$759,992**
Meta**0%****$484,093**
Google**0%****$376,445**
SpaceXAI**0%****$351,851**
DeepSeek**0%****$187,167**

Those displayed zeroes should be read as rounded market prices, not proof that an outcome is literally impossible. Prediction-market percentages reflect what traders will currently pay, subject to contract rules, liquidity, timing, and the information available before settlement.

The timing matters most. With only days remaining, the market is effectively pricing whether the present leaderboard order will survive until the designated snapshot. It is not forecasting the AI industry’s winner for 2027.

Prediction Bubbles 🫧 @predictionbubbl Sep 21, 2026

@Polymarket puts @AnthropicAI at 98.6% for best AI model on 30 September. @Kalshi has Claude at 72.5% for 31 December.

Ten days out the market is nearly certain. Three months out it is not.

@OpenAI is 0.4% for September and 13.4% for December.

https://predictionbubbles.net

View on X

That comparison captures the market’s real message. Polymarket reportedly put Anthropic at roughly 98.6% for September, while a December Kalshi market priced Claude at 72.5% and OpenAI at 13.4%. Near-term certainty coexists with substantial medium-term uncertainty. Public prediction-market coverage similarly tracks these contracts as fast-moving expectations rather than durable product rankings.[2][3]

BG-VC @bgvc123 Sep 15, 2026

Polymarket gives Anthropic a 91% shot at September’s “best” AI model. Sounds decisive—until the fine print: one text leaderboard settles it. Nearly $500,000 can price conviction. It can’t tell you which AI is best for your work. “Best” may be the laziest question in AI.

View on X

For practitioners, the useful interpretation is: traders believe Anthropic has enough of a current lead, under the settlement rules, that there is little time for a rival to reverse it. That is much narrower than declaring the AI race over.

Why does one leaderboard make “best AI model” a loaded question?

The contract resolves using a specified Arena.ai text leaderboard snapshot rather than a composite of production reliability, coding cost, security, latency, tool use, or enterprise adoption.[1] That gives traders an objective settlement mechanism, but it also compresses a multidimensional buying decision into one rank.

A leaderboard is useful when its evaluation resembles your work. It becomes misleading when teams treat the resulting rank as a universal recommendation.

A customer-support company may prioritize low latency, predictable structured output, and cheap high-volume inference. A coding-agent vendor may care more about completing long, multi-file tasks without human intervention. A regulated enterprise may accept lower benchmark performance in exchange for stronger governance, deployment controls, and contractual assurances.

RM (Polymarket Data) @resolvedmarket Sep 24, 2026

BREAKING: Anthropic's Claude Opus 5.5 knocked OpenAI off the top of https://Arena.ai Code Arena WebDev leaderboard.

Anthropic launched Opus 5.5 last week, undercutting its predecessor's price by 20%.

Polymarket's odds of OpenAI settling for second-best jumped from 23% to 90% overnight.

$100 on YES turned into $234.

Anthropic's lead is just 26 leaderboard points, thin enough to flip again before month-end.

View on X

The 26-point gap discussed in the live X conversation illustrates how a thin numerical lead can still support a 99% price. Probability is not the same thing as margin. If the remaining window is short and traders see no scheduled release likely to change the ranking, even a narrow lead can look highly defensible.

The reverse is also true: a small unexpected release, leaderboard correction, or model update can reprice the contract quickly. Earlier reporting on the market likewise showed odds changing as new models and rankings arrived.[4][5]

Developers should therefore separate two questions:

  1. Which company is most likely to occupy the contract’s chosen leaderboard position on September 30?
  2. Which model produces the best quality, speed, reliability, and cost on our workload?

Polymarket is designed to answer the first. Procurement requires the second.

Why are traders positioned so heavily behind Anthropic?

The simplest explanation is that Anthropic shipped a model at the right time with the right benchmark profile.

Claude Opus 5.5 launched on September 22, 2026, only days before the contract’s resolution window.[11] Anthropic’s published materials position it as a frontier model for coding, agents, and complex knowledge work.[7] The benchmark table discussed across X gives it 66.4% on Terminal-Bench 4.0, while launch reporting described substantially lower costs than Opus 5.[7][10]

st1ne @SolSt1ne Sep 22, 2026

Anthropic just shipped Claude Opus 5.5 and didn't bother with a press tour - just dropped a benchmark table against everything else on the market

The shift here isn't a new modality or a new trick. It's the same model, still faster, still cheaper, now just further ahead on the numbers that actually predict whether an agent finishes the job

here's the breakdown:

1 - agentic coding, three ways → 66.4% on Terminal-Bench 4.0 (Astra 57.9%, Fable 5.1 55.8%, Sol 37.3%), 54.4% on FrontierCode v1.1, 57.8% on CursorBench 4.0. Not a one-benchmark fluke - it leads on every coding harness that got tested

2 - knowledge work → 1846 on GDPval-AA v2.1, ahead of Fable 5.1 (1735), Opus 5 (1708), and well clear of GPT-6 Astra (1542) and Sol (1588)

3 - reasoning → 67.7% on Humanity's Last Exam with tools, the top score on the table, edging out Fable 5.1 (65.6%) and beating Astra by 10.5 points (57.2%)

4 - computer use → 81.8% on OSWorld 2.0 (partial), ahead of Fable 5.1 at 80.7% - Astra didn't even report a number here

5 - the honest catch → it's not a clean sweep either. Astra actually wins two categories outright: AutomationBench business workflows (41.4% vs Opus 5.5's 40.0%) and Terminal-Bench-Science (64.6% vs 58.7%). Worth knowing before you pick a model by vibes

6 - the stack worth building → pair Opus 5.5 with Jev (TypeSafe AI's new "System One" model, out Sept 15). Jev doesn't write text - it takes unstructured input and returns a typed decision with a calibrated probability in 70-500ms, at a fraction of a cent per call. So Opus 5.5 does the actual thinking - the agentic coding, the long research runs, the judgment calls - and Jev sits inside the loop doing the thousand small classifications and routing decisions the agent doesn't need a full model turn for. System 2 for the hard parts, System 1 for everything that repeats

7 - why this matters → most teams are running one model for both jobs right now - reasoning through a plan AND deciding "is this ticket urgent, yes or no" with the same expensive call. That's the gap Jev is built for, and it's the gap Opus 5.5's own numbers make obvious: the model is good enough that burning it on trivial decisions is waste

Full benchmark table and the Jev docs are worth five minutes before you touch your stack

View on X

That combination matters more than a raw intelligence claim. A model that completes more agentic work while consuming less budget can improve both application quality and gross margin. It also offers traders a visible catalyst immediately before settlement.

The same discussion shows why no benchmark table should be read selectively. The post notes categories in which Astra leads, including business-workflow automation and scientific terminal tasks. Anthropic may have the stronger overall position for this contract while OpenAI remains better suited to particular workloads.

*Walter Bloomberg @DeItaone Aug 17, 2026

CLAUDE TOPS AI RANKINGS AS COSTS FALL

Bank of America launched its Frontier AI Tracker, monitoring model intelligence, usage, token prices and hardware costs.

Anthropic’s Claude Opus 5 ranks #1 for intelligence, followed by Claude Fable 5 and OpenAI’s GPT-5.6 Sol.

DeepSeek leads usage share at 30%, while Anthropic dominates AI spending at 65%.

Meanwhile, AI token prices fell 9% month-over-month, helped by major OpenAI price cuts, while GPU rental costs remained broadly stable.

View on X

The Bank of America tracker figures circulating on X sharpen that point: Claude was described as leading intelligence rankings, DeepSeek as leading usage share, and Anthropic as taking the largest spending share. Those are three distinct measures—capability, consumption, and revenue—and they need not identify the same winner.

For September traders, however, better leaderboard performance at lower reported cost is precisely the kind of evidence that supports a near-term Anthropic position. For buyers, it is a reason to run an evaluation, not to skip one.

Could OpenAI’s effort scaling and inference economics overturn the consensus later?

The strongest counterargument is not that traders have misread September. It is that they may be extrapolating too much from a model optimized for the current evaluation regime.

Frank Downing @downingARK 2026-09-22T19:31:45.000Z

There is something clearly different in how Anthropic & OpenAI scale effort levels. Doesn't show up on every benchmark, but these results on FrontierCode make it really clear.

Anthropic models peak at lower effort levels, where as OpenAI models start low and climb up fairly consistently.

The result is a better score at a lower cost for Anthropic models, but an unintuitive experience where increasing effort does not increase scores and might actually degrade performance (on this benchmark at least).

View on X

The FrontierCode pattern described by Frank Downing suggests Anthropic and OpenAI models respond differently when given more inference effort—the extra computation spent generating, checking, or refining an answer. Anthropic reportedly performs strongly at lower effort, while OpenAI starts lower and improves more consistently as effort increases.

That distinction creates different production profiles:

Lisan al Gaib @scaling01 2026-09-03T18:56:28.000Z

From April after the Mythos Preview and GPT-5.5 launches:
- "I think my scenario of OpenAI being ahead by 1-3 months by end of year is more likely after this launch"
- "it's quite likely now that OpenAI will out-accelerate Anthropic"

Anthropic obviously has a better model internally, but I don't think it's that much better and OpenAI just continued RL training for GPT-6.1-Astra

OpenAI is just starting to hill-climb, while Anthropic already had ~7 month of hillclimbing with a massive model

View on X

The OpenAI bull case is that continued reinforcement learning and inference optimization could produce faster improvement than the September snapshot implies. Reporting and benchmark trackers show how quickly launches can change comparative positions, while analysis of OpenAI launch periods indicates that model news does not always move prediction markets as expected.[6][12][13]

alice @aliceisplaying 2026-09-05T10:18:25.000Z

anthropic has only ~20% less compute than openai and yes fable is definitely big but i think openai is just way, way better at optimizing inference for some reason

View on X

The claim that OpenAI has roughly 20% more compute is part of the X debate, not a verified conclusion established by this contract. But it points toward the relevant strategic question: which company can turn compute into useful inference most efficiently?

That helps explain why September can look settled while December remains open. The market implies that the existing ranking is unlikely to flip within days; traders appear less confident that Anthropic can preserve it across another model-training and release cycle.

Why can leaderboard leadership and enterprise spending point to different winners?

A benchmark measures model behavior under defined tests. Enterprise-spend data measures purchasing inside organizations with existing contracts, integrations, compliance requirements, and switching costs.

Ara Kharazian @arakharazian 2026-09-15T16:27:57.000Z

OpenAI is winning enterprise spend at the frontier. As of this week, Astra takes 13% of enterprise AI spend vs. Fable (8%) per Ramp data.

Some early thoughts:
1. Anthropic took a big risk in its recent call to pace the frontier. It's frontier model has already fallen behind on adoption.

2. OpenAI's growth is primarily coming from shifts from Sol and some Anthropic models, net-new usage too. That's good for them and suggests some pricing power remains by having a good, competitive frontier model.

View on X

The Ramp figures cited in the X conversation put OpenAI’s Astra at 13% of enterprise AI spend and Anthropic’s Fable at 8%, despite Anthropic’s stronger position in the September prediction market. Even if those measures use different product and customer scopes, the divergence is strategically important: model leadership and commercial distribution price separately.

BetG8 @BetGeight Sep 12, 2026

Shoppers pay 21.6% more, but the leaderboard hasn't moved for Google: on Polymarket's best AI model by end of September it's at 2.1 cents. Anthropic ran to 96.5 today, up 8 points; OpenAI dropped 8 to 0.7. Pricing power and model leadership price separately.

View on X

A buyer may spend more with a provider because it offers preferred APIs, volume discounts, administrative controls, regional availability, or an easier path through procurement. None of those factors necessarily raises an Arena score.

Usage also differs from spend. The Bank of America figures discussed on X put DeepSeek near 30% usage share while Anthropic captured around 65% of spending. The same post said token prices fell 9% month over month while GPU rental costs stayed broadly stable. That suggests high usage does not automatically create pricing power, and high spend does not necessarily mean the largest request volume.

For founders, these are separate dashboards:

A 99% settlement probability can coexist with commercial vulnerability if leadership is expensive to serve or difficult to distribute.

What hidden cost-to-serve risk sits underneath Anthropic’s 99% odds?

The bear case against Anthropic is economic rather than benchmark-driven.

Brandon Gell @bran_don_gell Apr 7, 2026

Prediction: Claude has massively taken the lead right now because they offer a better product, but that comes at a massive cost.

Buyers have not realized that included in a Claude subscription is not enough tokens to get real work done and that overages will cost $400 to $1,000 per day per user. Anthropic will need to buy significantly more compute, but because they don't own their own data centers, the cost to serve will continue to go up.

Spend will shift gradually and then quickly back to OpenAI, who can offer comparable models but at a much lower cost basis because they own their own data centers. Cost of inference will become the only competitive advantage making this market a race to the bottom.

Apple or Google will buy or merge(!!!) with Anthropic.

View on X

Brandon Gell’s post advances an aggressive prediction: subscription allowances may prove insufficient for intensive work, overages could become extremely expensive, and Anthropic’s dependence on external compute could weaken its cost position. The proposed $400 to $1,000 per user per day is a prediction from the post, not a documented average customer bill, and buyers should not treat it as established pricing.

Still, the underlying risk is legitimate. Agentic applications can generate long contexts, repeated tool calls, retries, evaluations, and large outputs. A workflow that looks affordable in a short demo may become expensive when it runs continuously for thousands of users.

Anthropic’s lower reported Opus 5.5 cost helps address that pressure.[7][11] But falling market-wide token prices can simultaneously benefit customers and squeeze providers. If GPU costs do not fall at the same rate, frontier-model vendors must improve utilization, model efficiency, caching, routing, or infrastructure terms to protect margins.

SaaS buyers should therefore ask vendors for:

The September odds imply confidence in rank—not confidence in Anthropic’s long-term unit economics.

Is the real competition shifting from models to an AI operating layer?

The most consequential point in the X conversation is that the model may no longer be the full product.

Rishi @RishiUvaach Aug 28, 2026

Most people are still asking:

Opus or Sonnet?

That may already be the wrong question.

Claude is starting to look less like an AI model and more like an AI operating layer.

The models are only one part of it.

Around them, Anthropic is building the pieces required to move AI from a chat window into real software:

→ Reasoning models
→ MCP and tool connectivity
→ Agent frameworks and SDKs
→ APIs and managed agents
→ Memory and context systems
→ Security and permissions
→ Evaluation pipelines
→ Governance and observability
→ Production deployment infrastructure

And that changes what it means to be good at AI engineering.

Prompting is becoming table stakes.

The harder skill is understanding how an agent gets context, remembers what matters, accesses tools safely, evaluates its own output, survives failures, and operates reliably inside a production system.

That is a very different skillset from simply knowing which model tops a benchmark.

In 2026, the advantage may not belong to the person who knows the best model.

It may belong to the person who understands how the entire AI stack fits together.

Claude’s evolution is a good preview of where AI engineering itself is heading.

View on X

MCP, or Model Context Protocol, gives agents a standardized way to connect with tools and data. Around that sit software development kits, memory systems, permissions, evaluations, observability, and deployment infrastructure. Together, these components determine whether an agent can operate reliably rather than merely produce an impressive answer.

That is why agentic benchmarks such as Terminal-Bench can matter more to software teams than a general intelligence score. They attempt to measure whether a system completes a multi-step job in an environment where planning, tool use, error recovery, and persistence all matter.

Anthropic’s Opus product and system documentation emphasize this broader agentic and safety context, including evaluations that extend beyond simple chat quality.[8][9] Live benchmark sites also show why teams increasingly need multiple views rather than one universal table.[14][15]

Pascual ⚡ @0xPascual 2026-09-25T09:44:32.000Z

Two frontier models shipped today. Opus 5.5 and GPT-6 Sol. Both got the same prompt: animate a looping 3D pelican on a bicycle in Blender.

The timeline flooded with benchmark screenshots.

The interesting part is upstream of the render. The test ran bpy automation straight through API endpoints at $10 per million input tokens and $50 per million output. No manual rigging. No traditional asset pipeline. The script was the workflow.

That's the thing undermining studio labor costs nobody's talking about.

Meanwhile Polymarket has 'Will any other model be the best AI model on October 5, 2026?' at 67% - the field is still wide open, even after today's drops. 30% on claude-fable-5.1-max, but 67% on any other model. The market isn't convinced the frontier is settled.

View on X

The Blender example makes the platform shift concrete. The economically important feature is not only which model generated the best-looking pelican. It is that an API-driven script could become the asset-production workflow itself.

That favors providers that can become embedded operating layers. Once a company has built tool integrations, permission policies, evaluation suites, memory, and failure handling around one ecosystem, switching models becomes more expensive—even if another provider temporarily takes the leaderboard lead.

The market contract misses this lock-in because it settles on output preference. The SaaS market will increasingly settle on workflow control.

Can political and safety narratives eventually move Anthropic’s odds?

Anthropic occupies an unusual position: it is presented as both an aggressive frontier competitor and a company unusually vocal about advanced-AI risks.

NabiYok🍌 @nabi_sarvi Sep 24, 2026

the white house just labeled dario amodei the face of AI doomerism and on @polymarket that same lab is still sitting at 99% to have the best model when september ends.

they even priced an 8% shot that washington takes a stake in anthropic while the guy is about to brief the un on risk. china is out here 90% bullish on the whole thing and the crowd still treats claude like the lock. weird week to be the pessimist everyone keeps betting on.

View on X

The White House and Washington claims in this post are the poster’s characterization of the political narrative and separate prediction-market pricing, not conclusions established by the September model contract. But the tension is worth watching. Traders can believe Anthropic is likely to win a leaderboard snapshot while remaining uncertain about regulation, government relationships, or access to compute.

MarketCalled @MarketCalled Sep 10, 2026

OpenAI sits at 9 percent to hold the best AI model on September 30, flat on the day with GPT-Live-1 shipping. Anthropic holds 88. There is $548K on the OpenAI leg of that Polymarket event, so the price reads an API release as distinct from a frontier one.

View on X

Earlier in the market, traders apparently treated an API-oriented OpenAI release as different from a frontier-model event. That shows narrative discipline: not every product launch changes the variable the contract measures.

Anthropic’s system card documents extensive safety evaluations alongside capability claims.[9] Over time, regulation, deployment restrictions, or government procurement could affect adoption and pricing. With days left in September, however, the market implies that those narratives are unlikely to displace the current leaderboard signal before settlement.

What should developers, founders, and SaaS buyers do with these odds?

The right response depends on what decision you are making.

Developers: Pick the model that wins your workload

Use Opus 5.5 as a serious candidate for long-running coding, complex software engineering, and knowledge work, especially where its reported lower-cost performance maps to the task.[7]

NonTechOps @NonTechOps Sep 23, 2026

Today's Best AI model

Yesterday Anthropic Launched their newest model

Claude Opus 5.5 is best for long-running agentic coding workflows, complex multi-file software engineering, and knowledge-work tasks, outperforming GPT-6 Astra on coding efficiency, costing roughly 60% less

View on X

But build a representative evaluation set first. Measure completion rate, latency, retries, token consumption, structured-output validity, tool failures, and human correction time. A model can top a benchmark and still lose on your repository, language, customer data, or latency target.

Founders: Preserve the ability to route across providers

The gap between September and December odds is a warning against deep single-model dependency. Put model access behind an internal abstraction, maintain regression tests, and keep prompts and tool definitions portable where possible.

This matters most for early-stage SaaS companies whose gross margin depends directly on inference pricing. Use expensive frontier models for difficult reasoning, and route classification, extraction, or repetitive decisions to cheaper systems when quality permits.

Enterprise buyers: Contract for total cost, not list price

Ask for scenario-based estimates covering normal usage, long-context agents, retries, and peak periods. Evaluate governance, observability, data handling, and vendor switching alongside benchmark quality.

OpenAI may fit organizations prioritizing broad enterprise adoption and effort-scaled inference. Anthropic may fit teams prioritizing the current agentic-coding profile and lower reported cost for high-end work. Google, Meta, DeepSeek, and others remain relevant where ecosystem integration, deployment control, open weights, geography, or price matters more than this specific leaderboard.

Everyone: Treat 99% as a narrow signal

Polymarket currently implies that Anthropic is highly likely to satisfy one contract’s definition of “best” on September 30, 2026.[1] It does not imply that the market has solved model procurement.

The more durable signal is the difference between the September and December prices. Traders see a near-term leader, not a settled industry. The companies most likely to convert temporary model leadership into lasting SaaS power will be those that combine strong models with efficient inference, enterprise distribution, tool connectivity, governance, and control of the operating layer around agents.

Sources

[1] Which company has the best AI model end of September? — Polymarket

[2] AI Predictions & Real-Time Odds — Polymarket

[3] Which Company Has the Best AI Model in September 2026? Winner Odds — Lines.com

[4] Which company has the best AI model end of September Odds & Prediction Market Analysis — CryptoSlate

[5] Odds On: Which company will be able to claim best AI model in September? — Markets Insider

[6] OpenAI Blitzed. The Money Didn’t Move. — ComputeLeap

[7] Claude Opus 5.5 — Anthropic

[8] Claude Opus — Anthropic

[9] System Card: Claude Opus 5.5 — Anthropic

[10] Anthropic launches Claude Opus 5.5: Benchmarks, pricing, safety — Mashable

[11] Anthropic unveils Claude Opus 5.5 — Reuters

[12] Best Anthropic Models, September 2026 — BenchLM.ai

[13] LLM Leaderboard & AI Model Benchmarks, September 2026 — BenchLM.ai

[14] Who Is Winning the AI Race? Monthly LLM Leader Timeline — BenchLM.ai

[15] Live AI Benchmarks — MultipleChat