market-watch

Traders Are Betting 94% on Anthropic: What Polymarket's AI Model Odds Reveal About Where SaaS Is Heading in 2026

Polymarket odds put Anthropic at 94% for best AI model by end of September 2026. Analyze what these implied probabilities mean for developers and SaaS buyers. Discover the strategy behind the numbers.

👤 📅 August 31, 2026 ⏱️ 22 min read
AdTools Monster Mascot reviewing products: Traders Are Betting 94% on Anthropic: What Polymarket's AI M
How we research: This guide is compiled by the AdTools team from the linked sources below and current public discussion. Pricing and features change often, so please verify time-sensitive details with each vendor before making a decision.

The practical question for developers, founders, and SaaS buyers is not whether Anthropic is certain to have the best AI model at the end of September 2026. It is whether Polymarket’s overwhelming preference for Anthropic reveals a durable industry shift—or merely a crowded bet on today’s benchmark leader.

As of August 31, 2026, traders have put roughly $772,665 into Polymarket’s “Which company has the best AI model end of September?” market, scheduled to resolve around October 1 using an arena leaderboard. The market implies a 94% probability for Anthropic, versus 4% for OpenAI, 2% for Google, 1% for SpaceXAI, and rounded 0% prices for Alibaba and Z.ai.[7]

Bottom line: Traders currently price Anthropic as the overwhelming favorite to lead the specified model leaderboard at the end of September. That is a strong signal about near-term benchmark expectations, especially for coding and agentic workloads—but it is not proof that Anthropic will offer the best economics, developer platform, availability, or long-term competitive position.

The 94% signal: What exactly is Polymarket’s AI model bet saying?

The market’s quoted probabilities and traded amounts are:

CompanyImplied probabilityAmount traded
**Anthropic****94%****$104,393**
**OpenAI****4%****$53,748**
**Google****2%****$64,113**
**SpaceXAI****1%****$54,393**
**Alibaba****0%****$56,918**
**Z.ai****0%****$49,620**

These displayed percentages are market prices, not scientifically calibrated forecasts. Because the outcomes trade through separate contracts and prices are rounded, they need not add neatly to 100%. A displayed 0% should also be read as “priced very close to zero,” not as mathematical impossibility. The market’s exact resolution rules matter more than the headline wording.[7]

The striking feature is the difference between concentrated odds and distributed trading activity. Anthropic has attracted the most volume among the listed contracts, but traders have also committed meaningful amounts to companies priced as remote possibilities. Volume shows how much a contract has changed hands; it does not directly measure net confidence.

Still, 94% is an unusually concentrated market expectation. It effectively says traders see the late-September window as too short for the current hierarchy to change.

The All-In Podcast @theallinpod Dec 3, 2025

Jason's AI Pair Trade: Short OpenAI. Long Google, xAI, and Anthropic.

Why? OpenAI's competition is fierce.

"They're facing a Google firing on all cylinders, Anthropic, and Grok beating them in the leaderboards pretty consistently."

Polymarket has Google's Gemini 3 at ~87% to finish 2025 as the top-ranked LLM.

Over the last six months, Gemini has started to shrink ChatGPT's massive lead in traffic share.

View on X

That interpretation matches a recurring X pattern: prediction-market prices are increasingly being treated as a live scoreboard for the frontier labs. Earlier markets generated pair trades built around weakening confidence in OpenAI and rising confidence in Google, Anthropic, and xAI.

Rihard Jarc @RihardJarc Jun 20, 2025

Over the last few days, on Polymarket, $GOOGL's lead as the best AI model by the end of 2025 has increased even further.

Anthropic odds have also risen, while those of OpenAI and xAI have decreased.

While $GOOGL's market share in LLM search will be much lower than the market share it has on traditional search on the enterprise side if $GOOGL turns out to be the best model provider and on top of it offers them via GCP on their TPU infrastructure, GCP's value could be much more than the market current anticipates.

View on X

But those discussions also expose the limitation. A market designed to identify one leaderboard winner compresses several different competitions—model quality, infrastructure economics, consumer distribution, enterprise revenue and developer adoption—into one price.

Why does the market imply a 94% chance for Anthropic?

The simplest explanation is that traders are extrapolating from Anthropic’s recent model momentum. Claude Opus 5 launched on July 24, 2026, with Anthropic positioning it around advanced coding, agents and sustained enterprise workflows.[1] VentureBeat similarly described the release as a cheaper model aimed at coding, agentic systems and business deployment.[5]

That positioning closely matches what a leaderboard-based resolution is likely to reward. Coding and agentic evaluations test whether a model can plan, call tools, modify repositories and complete multistep work—not merely produce fluent chatbot responses.

The benchmark narrative is reinforced by the Bank of America Frontier AI Tracker figures circulating on X: Claude Opus 5 ranked first for intelligence, followed by Claude Fable 5 and OpenAI’s GPT-5.6 Sol. The same summary said Anthropic represented 65% of tracked AI spending, while DeepSeek led usage share at 30%.

Walter Bloomberg @DeItaone Aug 17, 2026

CLAUDE TOPS AI RANKINGS AS COSTS FALL

Bank of America launched its Frontier AI Tracker, monitoring model intelligence, usage, token prices and hardware costs.

Anthropic’s Claude Opus 5 ranks #1 for intelligence, followed by Claude Fable 5 and OpenAI’s GPT-5.6 Sol.

DeepSeek leads usage share at 30%, while Anthropic dominates AI spending at 65%.

Meanwhile, AI token prices fell 9% month-over-month, helped by major OpenAI price cuts, while GPU rental costs remained broadly stable.

View on X

The distinction between spending share and usage share is crucial. High usage can come from free tiers, low-cost inference or open deployment. High spending suggests customers are willing to pay for capability—and enterprise customers often care disproportionately about coding, tool use and workflow automation.

Anthropic’s reported revenue mix strengthens that thesis. Aakash Gupta’s summary of the company’s trajectory describes an approximately $9 billion revenue run rate, more than 300,000 business accounts, 80% of revenue from enterprise API customers and Claude Code above $1 billion in annualized revenue.

Aakash Gupta @aakashgupta Jan 7, 2026

Anthropic is going parabolic.

It just went from $183B to $350B in four months. That’s a 91% jump.

Their revenue run rate hit $9B by end of 2025. At $350B pre-money, they’re trading at roughly 39x ARR.

Meanwhile… OpenAI is at $500B (maybe seeking $750B) with ~$13B in ARR. That’s similar multiples!

But the revenue is totally different.

Anthropic gets 80% of revenue from enterprise API customers. More than 300,000 business accounts. Claude Code alone crossed $1B in annualized revenue.

OpenAI gets most of its revenue from ChatGPT’s 800+ million weekly users. Consumer-heavy.

So you’ve got two companies with nearly identical valuation multiples but COMPLETELY different revenue structures.

Anthropic projects $20-26B ARR for 2026. OpenAI projecting $20B by year end.

And Anthropic says they’ll be cash flow positive by 2028. OpenAI is projecting $14B in losses for 2026 and won’t turn profitable until 2029 or 2030.

Similar multiples now despite completely different short-term trajectories.

View on X

Those numbers come from the public industry conversation rather than the prediction market’s resolution criteria. Even so, they help explain trader positioning: the market implies that Anthropic’s benchmark lead is supported by a commercial feedback loop involving paid API usage, enterprise workflows and coding products.

For teams, that makes Anthropic the market’s apparent default when frontier coding quality is the first requirement and premium API pricing is acceptable. It does not automatically make it the right choice for consumer-scale applications or cost-sensitive inference.

Could OpenAI’s 4% probability be underpriced?

OpenAI’s 4% implied probability looks dismissive until the time structure is considered. The market resolves on one leaderboard snapshot around October 1. A competitive model shipped during September could change rankings—and contract prices—very quickly.

At 4%, traders are mostly betting that no OpenAI release will displace Anthropic within the remaining window. They are not pricing OpenAI as having only a 4% chance of remaining an important AI company.

The contrarian argument is that OpenAI has historically been more willing to ship frontier capability rapidly, including models with consequential cyber capabilities, while Anthropic has used a more cautious release posture.

Yuchen Jin @Yuchenj_UW Aug 7, 2026

OpenAI and Anthropic are taking very different strategies.

OpenAI seems willing to ship frontier models with cyber capabilities to the public asap, Anthropic is more cautious.

That means Astra could very well be the best model in the world when it launches.

For the first time in a long time, GPT will be ahead of Claude.

View on X

If that assessment is right, a new OpenAI model could create asymmetric movement: limited downside for a contract already priced near the floor, but substantial upside if a release performs well on the designated arena. That is a market argument, not a prediction that such a launch will occur.

The longer-term bull case focuses on capacity. One X thesis argues that OpenAI’s earlier compute commitments and custom-ASIC ambitions could eventually reduce inference costs, increase usage allowances and strengthen its position as a platform.

Equity Climb @equityclimb1 Aug 27, 2026

I think OpenAI has played the long game much better than Anthropic, and over time that advantage could become difficult to overcome.

OpenAI aggressively secured compute last year, and that bet is now paying off. Anthropic took a more cautious approach, and today Claude is dealing with tight usage limits, capacity constraints, and frequent outages.

The real wildcard is OpenAI’s custom ASIC. If it reaches mass production and scales well, OpenAI could dramatically lower inference costs, offer far more generous usage, and attract even more users and developers.

Sam has repeatedly described OpenAI as a platform that other companies will build on. Anthropic seems to be moving in the opposite direction—pushing deeper into downstream applications like coding and biology.

One wants to become the foundation of the ecosystem. The other risks competing with the companies it needs to build on top of it.

View on X

That would not necessarily affect the September result. It matters because model competitions increasingly have two clocks:

  1. The release clock, which determines who tops a benchmark this month.
  2. The infrastructure clock, which determines who can serve strong models cheaply and reliably for years.

Prediction markets can adjust quickly after a launch, but before that launch they often represent a weighted vote for the status quo. OpenAI’s 4% contract is therefore best understood as a low-probability bet on a near-term ranking change—not a valuation of its entire strategy.

Is compute ownership the hidden variable behind the odds?

The September market asks about model performance, but the durable competition may be decided by the cost of producing each useful token.

Google owns extensive data-center infrastructure and designs TPUs. xAI has pursued direct infrastructure build-out. OpenAI and Anthropic rely heavily on hyperscaler relationships for compute, with those cloud providers also serving as strategic investors or partners. The sharp version of the argument on X is that renting versus owning compute produces a structural gap once models become harder to differentiate.

Jun Song @jun_song Aug 14, 2026

The difference between OpenAI/Anthropic and xAI/Google:

1. OpenAI and Anthropic rent compute from hyperscalers. xAI and Google own their data centers.
2. Most of OpenAI and Anthropic's equity comes from hyperscaler capital. The other two aren't built on that model.

This creates a massive gap in price competition.

Up until now, they justified high prices with superior performance. But that performance gap has narrowed, and because of those structural costs, they can't offer competitive pricing anymore.

Unless they pull ahead again with a massive breakthrough, OpenAI and Anthropic will eventually get acquired by their investors, Amazon and Microsoft.

View on X

That matters because pricing pressure is intensifying. The Bank of America tracker summary reported a 9% month-over-month decline in AI token prices, while GPU rental costs remained broadly stable. If selling prices fall faster than underlying compute costs, margins tighten—even for the model leader.

The market’s 2% implied probability for Google may therefore understate Google’s broader strategic position. A leaderboard can say little about the value of combining competitive models with TPUs, Google Cloud, Workspace, Android, Search and an existing advertising engine. Earlier Polymarket discussions recognized that a leading Google model delivered through GCP on TPU infrastructure could have consequences beyond the model ranking itself.

For buyers, the decision criteria are different from the market’s:

The winning model contract and the winning gross-margin structure may point to different companies.

Does Anthropic’s lead validate B2B AI over consumer AI?

Anthropic’s 94% odds can also be read as a vote for an enterprise/API-first model strategy. Enterprise customers pay for models that can generate code, invoke tools and complete expensive knowledge work. Those are also the capabilities prominent benchmarks increasingly attempt to measure.

Consumer AI follows a different economic logic. It requires large-scale distribution, low-cost inference and, potentially, advertising. Google already possesses those advantages. OpenAI is trying to compete across both consumer and enterprise markets, creating the risk that neither receives complete strategic focus.

Rihard Jarc @RihardJarc Mar 23, 2026

It is becoming clearer every day that AI labs, as they transition from research organizations to "real" companies dependent on revenues and profits, will have to focus on either the enterprise or consumer path. Anthropic is clearly choosing the B2B path; $GOOGL is leaning heavily into B2C, while OpenAI wants to capture both, but in doing so risks losing the dominant position in either.

It seems the AI subscription/usage business model for enterprises is working well and has room to grow, but for consumer AI usage, the ad model will be key, and OpenAI is entering the arena where $GOOGL is the king.

Building a successful ad platform will be a challenge for OpenAI. Building out a good ad ecosystem at scale is much harder than people expect. On scale, $META and $GOOGL have really mastered it, while many other platforms have struggled for years.

View on X

This split explains why Google can be priced at only 2% in the September model market while remaining formidable in the wider AI business. Consumer success might come from embedding “good enough” models in products used by billions, not from topping a specific frontier leaderboard.

The inverse applies to Anthropic. Its market-favored model position does not mean it has won consumer distribution. Rather, traders currently imply that the enterprise capabilities Anthropic prioritizes are most likely to score well under this market’s resolution method.

For SaaS founders, the lesson is not simply “use Claude.” It is that B2B AI revenue appears increasingly tied to completing valuable workflows rather than maximizing chat engagement. Products selling to businesses should measure task completion, human-review time, failure rates and cost per successful workflow. Consumer products should focus more heavily on latency, retention, distribution and inference cost.

Is the “best model” becoming the wrong question?

A benchmark market produces a clear winner because it must. Production systems rarely do.

Developers increasingly buy an operating layer consisting of model access, tool connections, agent frameworks, memory, permissions, evaluation, observability and deployment controls. Rishi’s X framing captures the shift: Claude is beginning to look less like one model and more like a stack for operating AI inside software.

Rishi @RishiUvaach Aug 28, 2026

Most people are still asking:

Opus or Sonnet?

That may already be the wrong question.

Claude is starting to look less like an AI model and more like an AI operating layer.

The models are only one part of it.

Around them, Anthropic is building the pieces required to move AI from a chat window into real software:

→ Reasoning models
→ MCP and tool connectivity
→ Agent frameworks and SDKs
→ APIs and managed agents
→ Memory and context systems
→ Security and permissions
→ Evaluation pipelines
→ Governance and observability
→ Production deployment infrastructure

And that changes what it means to be good at AI engineering.

Prompting is becoming table stakes.

The harder skill is understanding how an agent gets context, remembers what matters, accesses tools safely, evaluates its own output, survives failures, and operates reliably inside a production system.

That is a very different skillset from simply knowing which model tops a benchmark.

In 2026, the advantage may not belong to the person who knows the best model.

It may belong to the person who understands how the entire AI stack fits together.

Claude’s evolution is a good preview of where AI engineering itself is heading.

View on X

This is why the market’s 94% signal is both meaningful and incomplete. Anthropic may be strongly favored under the chosen benchmark while a different provider wins a deployment based on:

Competition is also moving above the model layer. Developers can use harnesses that route tasks among multiple providers. The X conversation points to DeepSeek’s open, customizable Harness and to competing developer ecosystems from OpenAI and SpaceX/xAI.

Shruti @heyshrutimishra Aug 19, 2026

Claude is losing the AI war

While they're extending limits and asking for more time, their competitors aren't waiting

The first shot came from China. DeepSeek shipped Harness last week. You can self-host it, run whatever model you want, and customize everything through plugins. It's Claude Code for free. Over 160,000 GitHub stars in less than a week.

SpaceX is building their own developer ecosystem. They acquired Cursor and launched Origin yesterday, code hosting built for AI agents.

OpenAI Codex is growing fast. Developers are switching.

Kimi is becoming the cheaper alternative. And they're moving fast.

I myself used to love Claude, but now removing Claude from my stack. Their guardrails are killing the product.

View on X

The exact adoption claims in that post should be treated as part of the live practitioner debate, but the strategic direction is clear: model portability is becoming a product feature. As orchestration improves, the value of a marginal benchmark advantage may decline unless the provider can convert it into reliable workflow performance.

Anthropic’s official model catalog and independent model listings show how quickly model choices and API identifiers can proliferate.[4] Buyers should therefore evaluate the surrounding platform and migration path, not hard-code a product around one leaderboard leader.

Could usage limits and safety friction make the 94% lead fragile?

The strongest challenge to Anthropic’s market position is not necessarily a weaker benchmark score. It is whether developers can reliably access the capability being benchmarked.

Power users on X report assembling competitive stacks from multiple providers and canceling high-end Claude subscriptions. One former Claude Pro and Max subscriber argued that “harness freedom”—routing work among GLM, GPT and Kimi models—now provides a frontier experience without depending on Anthropic.

Mohamed Messaad @sshbeetle Aug 31, 2026

After almost 2 years of Claude Pro then Max 20x, I finally cancelled

It's now possible to have a frontier experience without Anthropic, which wasn't the case just a month ago

Current stack:
- @pidotdev with GLM 5.3 Flash, GPT 5.6 Sol as advisor through my Codex sub, Kimi K3 for front-end
- Codex with GPT 5.6 Sol for heavy-duty stuff (a bit faster than Pi when using only Sol)

Honestly haven't been missing anything. Fable 5 is great for front-end but not better than Kimi K3, for the rest GPT 5.6 Sol blows it out of the water and is almost unlimited with the resets

The future is in harness freedom

View on X

This is anecdotal, not representative churn data. But it identifies a genuine platform risk: if developers can substitute among models, tight usage limits and outages impose a higher penalty. A superior model that is unavailable during a production task can deliver less business value than a slightly weaker model with predictable capacity.

Safety behavior creates another tradeoff. Anthropic emphasizes cautious deployment, but agentic software must operate in hostile environments involving untrusted repositories, tools and dependencies. One widely discussed example involved Claude Code refusing an attacker-provided decoder, writing an alternative and still importing a poisoned module.

Kerem — road to $100k @mkeremturhan Aug 31, 2026

Claude Code refused to run the attacker's decoder, wrote its own instead, and the module it imported was poisoned.

Refusing was correct. It lost anyway.

Anthropic's own words for the feature: "best-effort classifier, not a security guarantee." Believe them and sandbox it.

View on X

The important conclusion is the poster’s own: a “best-effort classifier” is not a security guarantee, so agentic execution still requires sandboxing. Independent benchmark coverage of earlier Claude releases also highlights that intelligence, speed and price must be evaluated together rather than reduced to one ranking.[6]

These issues could reprice the market only if they affect the designated leaderboard or coincide with a competitor’s release before resolution. For SaaS buyers, however, they matter immediately. Teams should test rate limits, refusal patterns, outage handling and tool security before selecting a primary provider.

Are Alibaba and Z.ai really 0% threats?

Polymarket displays 0% implied probabilities for Alibaba and Z.ai, despite approximately $56,918 and $49,620 in traded volume, respectively. Again, these are rounded prices near zero—not declarations that either company has zero technical or commercial relevance.[7]

The Chinese-model discussion highlights the tension between arena position and market expectations.

Amelia @Ameliawang2014 Aug 25, 2026

Three numbers frame Aug 31's AI market: 📊

Qwen 11, https://chat.z.ai/ 15, Moonshot 17 on Arena's Aug 21 table. Alibaba is priced at 94.6% on Polymarket. I favor the leader—but the ranking can still move.

#Alibaba #Qwen #ChineseAI #Polymarket

View on X

Alibaba’s Qwen and Z.ai can have meaningful leaderboard presence without traders expecting them to finish first under the specific resolution criteria. DeepSeek’s reported lead in usage share makes the same point: adoption, especially at lower prices, is different from holding the top benchmark position.

The strongest counterargument is that the frontier remains concentrated among Anthropic and OpenAI, with other labs still materially behind despite rapid progress.

🍓🍓🍓 @iruletheworldmo Jun 29, 2026

i get why people want to root for “open source”.

but the distance between openai/anthropic and anything else is gargantuan. and it isn’t only open source that’s miles back, the other closed for-profits are too.

google, meta and xai are nowhere near. only two labs are sitting at the actual frontier, and the government keeps telling you which two: it force-pulled anthropic’s two best models overnight, and made openai submit its newest one to user screening before it would let it ship. it’s doing that to no one else, because there’s nothing else worth controlling.

and even if we only look at the publicly available models from these two, they dwarf anything held back privately by any company on the planet.

whilst mythos feels like another paradigm shift, it’s the result of pushing the scaling laws further than anyone else can. people misunderstand scaling as one single axis to push, when there’s so much left to scale across all of them: pre-training compute, post-training and rl, test-time compute, data.

you’ll start seeing mythos like jumps every two months, opus 4.7 to 4.8 was already about that and 5.5 to 5.6 runs on the same clock, as we’re now deep inside a hard, fast, and turbulent take off scenario.

so as all the best models say, buckle up buttercup.

View on X

Polymarket pricing largely endorses the first half of that thesis—but more strongly for Anthropic than OpenAI. Yet buyers should not translate near-zero winner odds into near-zero procurement value.

Alibaba, Z.ai and other open or lower-cost providers may fit when:

“Unlikely to rank first” and “not economically useful” are entirely different judgments.

What should developers, founders and SaaS buyers do with these odds?

The market’s clearest message is that traders expect Anthropic’s current model momentum to survive through the end of September. The wrong response is to convert that short-term probability into permanent vendor lock-in.

Developers: preserve model and harness freedom

Build an abstraction layer for prompts, tool schemas and outputs. Maintain eval sets that can be run against at least two providers. Sandbox agentic tools regardless of the vendor’s safety claims.

Also verify whether the bottleneck is really the model. Some practitioners argue that existing frontier models are already underused because teams lack context engineering, evaluation and workflow design.

Anatoli Kopadze @AnatoliKopadze Aug 31, 2026

Anthropic engineer:

"Opus 5 is already smarter than we know how to use. The bottleneck was never the model, it's you."

In 19 minutes he shows exactly how to get everything out of Claude with no extra tools, no extra costs.

You already pay for all of it and at best using 10% of what it can do.

Watch the session, then read the guide below on the Claude features 99% of users never find.

View on X

Founders: buy completed work, not benchmark prestige

Use model benchmarks to create a shortlist, then evaluate:

  1. Cost per successful task
  2. Tail latency
  3. Rate-limit behavior
  4. Human correction time
  5. Tool-call and structured-output reliability
  6. Security and compliance controls
  7. Failover difficulty

Anthropic fits teams monetizing premium enterprise automation. Google may fit companies already centered on GCP or serving consumer-scale workloads. OpenAI remains relevant where broad ecosystem reach and rapid shipping matter. Open and Chinese models deserve testing when cost or deployment control dominates.

SaaS buyers: treat 94% as a snapshot, not a procurement mandate

Negotiate data-export rights, portable prompt and evaluation assets, transparent usage limits and the ability to route overflow traffic. Watch September releases because a single strong launch could materially change the market before October 1.

Finally, cross-check prediction prices against model leaderboards and benchmark methodology. Coding rankings, general intelligence evaluations and arena preference tests measure different things.[13][14][15] Google, in particular, remains a strategic contender even when traders assign it only a 2% probability in this narrow contest.

Lisan al Gaib @scaling01 Mar 14, 2026

I would genuinely love for this to happen

but many people think that OpenAI and Anthropic are already in a positive feedback loop

and as we have seen with Gemini 3 Pro: a ~5 trillion param reasoning model won't magically be AGI
(or for that matter a 6T param Grok-5)

my base case is that OpenAI and Anthropic will pull further ahead

xAI has less compute, less researchers, less data (no Codex, no Claude Code) and does not have access to models that literally speed up research (behind ~6 months)

Google on the other hand is still in the race, being only ~3 months behind. they have the most compute, researchers, an infinite money glitch and the data

View on X

The deeper signal from Polymarket is not simply that traders favor Anthropic. It is that the industry is separating into model leaders, infrastructure leaders, distribution leaders and operating-layer leaders. One company may occupy several positions, but buyers should not assume the September leaderboard will identify all of them.

Sources

[1] Anthropic, “Introducing Claude Opus 5” — https://www.anthropic.com/news/claude-opus-5

[4] BenchLM.ai, “Claude Models & API IDs: Current Anthropic Model List” — https://benchlm.ai/providers/anthropic

[5] VentureBeat, “Anthropic launches Claude Opus 5, a cheaper AI model for coding, agents and enterprise workflows” — https://venturebeat.com/orchestration/anthropic-launches-claude-opus-5-a-cheaper-ai-model-for-coding-agents-and-enterprise-workflows

[6] Artificial Analysis, “Claude Opus 4.8 — The new #1 AI model” — https://artificialanalysis.ai/articles/claude-opus-4-8-analysis-and-benchmarks

[7] Polymarket, “Which company has the best AI model end of September?” — https://polymarket.com/event/which-company-has-the-best-ai-model-end-of-september-20260717143435868

[13] LM Market Cap, “Best AI LLM Models Ranked (2026)” — https://lmmarketcap.com/best/coding

[14] LM Market Cap, “AI Benchmarks 2026 — MMLU, GPQA, SWE-bench” — https://lmmarketcap.com/benchmarks

[15] SevenLab, “AI Leaderboard 2026: Top LLMs Ranked Daily” — https://www.sevenlab.ai/ai-leaderboard