The Best AI Model Odds in 2026: What Polymarket's $2M Bet Reveals
Polymarket odds put Anthropic at 54% for the best AI model of 2026, with Google at 36%. See what the smart money signals for developers and SaaS buyers.

The real question for developers, founders, and SaaS buyers is not simply which lab will win? It is whether prediction-market traders are identifying a durable technical leader—or merely betting on which model will occupy one leaderboard position on one date.
As of October 2, 2026, roughly $1,987,584 has been traded in Polymarket’s “Which company has best AI model end of 2026?” market, scheduled to resolve around December 31.[1] Traders currently price Anthropic at 54%, Google at 36%, OpenAI at 6%, xAI at 1%, DeepSeek at 0%, and Alibaba at 1%. Those prices make Anthropic the favorite and Google the only close challenger, but they do not establish either company as the future winner.
Bottom line: The market implies a roughly nine-in-ten chance that Anthropic or Google takes the year-end title. But fragmented benchmarks, a large disagreement with Kalshi, and the difference between leaderboard leadership and production value make these odds more useful as a sentiment gauge than as a technology-procurement guide.
The $2M signal: How are traders pricing the frontier race?
Prediction-market prices compress dispersed opinions into a probability-like number. A 54% Anthropic price means traders currently value its contract as though Anthropic has approximately a 54% chance of satisfying the market’s resolution rules. It does not mean Anthropic is 54% better, nor that a majority of developers prefer Claude.
The full October 2 snapshot is:
| Company | Implied probability | Traded volume |
|---|---|---|
| Anthropic | 54% | $279,909 |
| 36% | $181,767 | |
| OpenAI | 6% | $180,658 |
| xAI | 1% | $148,016 |
| DeepSeek | 0% | $125,201 |
| Alibaba | 1% | $124,578 |
The displayed percentages total 98%, which can occur because of rounding, spreads, and the mechanics of separate contracts. Volume is also cumulative trading activity, not necessarily capital currently at risk or evidence that every dollar represents an independent informed view.
Polymarket’s contract rules—not an abstract consensus about “intelligence”—ultimately determine resolution.[1] That makes the expected leaderboard or evaluation mechanism critically important. Traders may rationally buy the company most likely to win the designated ranking, even if they believe another vendor offers better image generation, lower inference costs, or a stronger enterprise platform.
The rapidly changing snapshots posted on X show how sensitive these odds are to releases and market definitions:
anthropic 52%, google 41%, openai 3% for best model by end of year
https://poly.market/jwYulOQ
anthropic 72%, openai 11%, google 10% to have the top AI model by end of 2026
https://polymarket.us/event/company1-2026-12-31?utm_source=twitter&utm_medium=post&utm_campaign=%40AskPolymarket
Prediction markets can aggregate release rumors, benchmark results, product access, and trader expectations faster than conventional analyst reports. But the conflicting snapshots also show their weakness: a price is only as robust as the contract language, liquidity, trader mix, and information available at that moment. Polymarket’s broader AI markets provide useful context, but none turns uncertainty into certainty.[7]
Why does the market favor Anthropic at 54%?
The market’s Anthropic position appears to price a specific thesis: coding and agentic software development are becoming the most commercially important frontier-model workloads, and Claude’s advantage there may persist through year-end.
That thesis is visible in the practitioner conversation:
Anthropic used more coding data in their training runs, so Claude is better at coding. DeepMind now knows this and sees agentic coding taking off exponentially in terms of revenue, so they will use more coding data for future Gemini models to catch up in capability.
"DeepMind engineers use Claude as a daily tool. Most of the rest of Google does not. When the question of equalizing access came up internally, the proposed response was to remove Claude for everyone — which DeepMind objected to so strongly that several engineers reportedly threatened to leave."
The claim that Anthropic trained more heavily on code is difficult to reduce to a single public metric. Yet the broader feedback-loop argument matters. A model that attracts serious developers generates more interactions around repository navigation, debugging, tool use, and long-running agent tasks. Those products can then expose failure patterns that inform evaluation and future development.
Current rankings also provide some support without producing a universal verdict. One widely circulated September summary puts Claude Opus 5.5 at roughly 58 on the Artificial Analysis Index, ahead of Gemini 4 Argon and GPT-6 Astra on that measure. Other composite rankings produce different winners, underscoring how evaluation selection changes the answer.[9][11]
Top frontier models by recent Sep 2026 benchmarks (composites, AA Index, Arena, coding/knowledge):
Claude Opus 5.5 (Anthropic) leads AA Index (~58)
GPT-6 Astra (OpenAI) tops many overall composites (BenchAlign ~88)
Gemini 4 Argon (Google) #1 Text Arena (1525), Vals Index, DeepSWE 77.9%; matches Astra on AA (53); strong long-horizon, cyber, multimodal
Claude Fable 5.1 / Sonnet 5.5
Others: Muse Spark 1.3, Grok 4.7
Rankings vary by task; Argon just launched with limited access.
Traders may also be overweighting coding because it has a clearer willingness to pay than many consumer chatbot tasks. A coding agent that completes migrations, resolves tickets, or shortens review cycles can be tied directly to labor cost and delivery speed. That makes coding leadership commercially legible—even if it is not synonymous with overall intelligence.
Still, Anthropic’s 54% price leaves substantial room for failure. Claude can be better overall for some practitioners while remaining more expensive or narrower in modalities:
Since mid 2025, OpenAI has been cheaper but worse than Anthropic. Recently OpenAI became better, but then with Opus 5.5, Anthropic is now the best overall.
ChatGPT has access to Reddit and Gemini has access to YouTube - neither of which Claude has. Also ChatGPT can do images which Claude doesn’t have. Otherwise Claude is better but more expensive.
Who fits Anthropic today? Teams whose bottleneck is complex coding, repository-scale reasoning, debugging, or autonomous software work have the strongest reason to accept a premium. Teams dominated by high-volume summarization, images, or price-sensitive customer interactions should not infer from the 54% contract that Claude is automatically their best economic choice.
Is Google’s 36% probability underpricing Gemini 4 Argon?
Google’s 36% implied probability makes it the clear challenger. The bull case strengthened after the September 30 release of Gemini 4 Argon, which was reported as competitive across coding, workflow automation, long context, video, and cyber-defense evaluations.[12]
The headline figures circulating on X include 77.9% on DeepSWE v1.1, 51.3% on AutomationBench, and 84.2% on long-context GraphWalks at 256K–1M tokens. But Argon reportedly trails on other difficult coding and scientific terminal benchmarks, while its verbose output can increase per-task cost.
Google released Gemini 4 Argon on 30 September 2026, claiming frontier performance across software engineering, enterprise knowledge work and cyber defense. Here's what the benchmarks actually show.
📊 The numbers
• DeepSWE v1.1 (coding): 77.9%, vs Claude Opus 5.5 74.2% and GPT-6 Astra 74.1%
• AutomationBench (workflows): 51.3%, 8.8 pts ahead of Opus 5.5 (42.5%)
• GraphWalks 256K–1M (long context): 84.2%, 12.4 pts ahead of Astra (71.8%)
• LVBench (long video): 91.7%, vs Astra 87.5%
• FrontierSWE v2 (hard coding): 55.0%, last of four. Astra leads with 65.5%
• Terminal-Bench Science: 57.6%, behind Opus 5.5 (63.3%) and Astra (68.1%)
• Cost per task (Artificial Analysis): $1.99 vs $0.72 for GPT-6.1 Sol and $3.26 for Astra, because Argon writes ~62K tokens per task
The most interesting signal may be reliability rather than raw accuracy. On AA-Omniscience, Argon was reported with a 15% hallucination rate, compared with higher rates for several competitors. Its accuracy was not necessarily highest; the model appears more willing to acknowledge uncertainty instead of inventing an answer.
Gemini 4 Argon has the lowest hallucination rate among leading models on Artificial Analysis' AA-Omniscience benchmark.
Gemini 4 Argon: 15%
Grok 4.7: 29%
GPT-6 Astra : 45%
Opus 5.5: 59%
This matters beyond model error counts. Intelligence isn't only getting answers right. It includes metacognition: knowing how reliable your own knowledge is. When a model's confidence metric tracks accuracy, uncertainty becomes a signal to retrieve, ask, or defer. A hallucinating model has a poor signal. So a lower rate suggests the model has some sense of where its knowledge ends, and that's what autonomy rests on. Argon is where the state of the art stands today.
That distinction matters in production. A model that knows when to retrieve information, ask a user, or escalate can be more useful for autonomous workflows than one with higher average accuracy but poorly calibrated confidence.
Google’s strategic argument is therefore broader than “Gemini wins coding.” Its strengths may fit document-heavy work, computer use, visual understanding, long conversations, and integration into existing enterprise tools:
A lot of people are saying Google is falling behind after Gemini 3.6 Flash.
I think they're reading it the wrong way.
To me, Google has changed its strategy.
Yes, Gemini is behind GPT-5.6 Luna, Grok 4.5, and Claude Sonnet 5 in coding.
But it leads in computer use, visual understanding, and long context. At 1 million tokens, it scores more than twice as high as Gemini 3.5 Flash.
That doesn't look like a company that is losing.
It looks like a company building for real work.
Frontier models are already smart enough for most thinking tasks.
Now the question is not who gets another benchmark record.
The question is who helps people do their jobs every day.
People on X talk about agents and hard benchmarks.
Most companies are still trying to figure out where AI fits.
Most workers are not running agent systems.
They need a model that can read documents, understand charts, keep track of long conversations, and work inside the tools they already use.
That is exactly where Gemini is strong.
If AGI is about doing every kind of knowledge work, then vision, long context, and real world understanding matter just as much as coding.
As Demis Hassabis has said, intelligence has to bring all of these things together.
Many developers think Google is losing the coding race.
I think Google has stopped chasing benchmark wins and started focusing on where the money is.
A fast, low-cost model that fits into everyday work may end up being the better strategy.
Near-term markets have reflected that possibility. One October contract reportedly put Google at 64% and Anthropic at 36%, suggesting traders believed Argon could hold a short-window lead even while the December market favored Anthropic.[2]
Anthropic will have the top-ranked AI model by the end of October, per Polymarket, at a 36% chance.
Google: 64%, Anthropic: 36%.
Can Gemini 4 Argon keep Google in the lead?
There is nevertheless a gap between benchmark availability and product availability. Argon’s limited initial access weakens its practical impact, and reported rankings still put Opus 5.5 ahead on the Artificial Analysis index.[11]
google's first non-Flash frontier model in over 7 months.
Artificial Analysis has Gemini 4 Argon (high) at 53 on the Intelligence Index. tied with GPT-6 Astra (max), 1 ahead of GPT-6.1 Sol (max). Opus 5.5 still sits at 58.
the weird AA number: 15% hallucination rate on AA-Omniscience vs 51% for Astra. accuracy is lower though (50% vs 63%). more "i don't know", fewer invented facts.
intro is $2/$10, then $4/$20. Fairwind cyber testers first. OpenRouter has no argon id this morning.
so google is back in the conversation. tied with Astra on paper, still behind Opus, and you mostly can't buy it yet.
Who fits Google? Organizations already committed to GCP, Workspace, multimodal document processing, or very long contexts have the strongest reason to watch Gemini. Google’s 36% price is especially relevant to buyers who value a combined model-cloud-productivity stack rather than a standalone API score.
Why do Polymarket and Kalshi disagree by roughly 20 points?
The sharpest warning against overconfidence is the cross-market gap. While Polymarket prices Anthropic near 54%, X market watchers report that Kalshi has put it near 74%—a disagreement of approximately 20 percentage points.
Huge 21% gap between prediction markets on the AI race.
Kalshi gives Anthropic a 74% chance of holding the top AI model by end of 2026. Polymarket is down at 53%.
Who has this right? 👇
https://polymarket.com/event/which-company-has-best-ai-model-end-of-2026?via=prediction-pantheon
That gap is too large to dismiss as ordinary noise. Several explanations are plausible:
- Different resolution rules: “Best model” can mean different leaderboards, dates, tie-breakers, or eligible releases.
- Different trader populations: One venue may attract more crypto-native traders; another may draw more US event-market participants.
- Different liquidity and spreads: A displayed price can move substantially when order books are thin.
- Different timing: AI odds can change immediately after a release, benchmark update, or access announcement.
- Different interpretations: Traders may disagree over whether a ranking measures raw capability, public availability, or an exact model variant.
Published market analyses citing Anthropic around 74%–75% may therefore reflect a different venue, timestamp, or contract interpretation rather than a contradiction in the underlying data.[2][3][4]
Shorter-duration markets can diverge too:
Same question, two prices. Anthropic having the top model at the end of October is 40 percent on Polymarket and 36 on Kalshi, a 4 point gap on $402K against $97K. Both rose today, 14 and 10, and 40 is the side that has to be wrong unless a new Claude ships this month.
View on XFor practitioners, this disagreement is more valuable than a falsely precise consensus. It says the outcome remains highly dependent on evaluation design and release timing. If sophisticated-enough markets cannot agree whether Anthropic’s probability is in the low-50s or mid-70s, a buyer should not treat either figure as a reason to lock an architecture to Claude.
Why are OpenAI at 6% and xAI at 1% despite strong products?
OpenAI’s 6% implied probability is the market’s biggest puzzle. GPT-6 Astra reportedly tops several composite evaluations even while Claude leads the AA Index and Gemini leads particular Arena and coding measures.[11] In other words, traders are not simply buying the company with the largest collection of benchmark wins.
They may instead be betting on the exact resolution mechanism, expecting another Anthropic release, or discounting OpenAI’s probability of holding the designated top spot on December 31. The price could also indicate that OpenAI’s commercial strategy increasingly optimizes for deployment breadth, speed, multimodality, and cost rather than one headline ranking.
That can still be attractive to enterprises:
I fully expect Grok & OpenAI to start taking enterprise marketshare away from Anthropic with their blend of intelligence, cost, and speed.
I haven't used Claude models in any meaningful sense since both Grok 4.5 and GPT 5.6 Sol dropped.
Fable 5 might be the most intelligent model, but its cost and speed make a non-starter for me, and I'm willing to bet the same for many people out there.
For SaaS companies, a slightly weaker model that responds faster and costs materially less can produce the better product. At scale, inference cost affects gross margin; latency affects conversion and user trust; multimodal support can eliminate the need for separate vendors.
xAI’s 1% price similarly refers to year-end ranking odds, not its ability to win particular workloads or customers. The bear case is structural: fewer researchers, less compute, and weaker coding-product feedback loops than OpenAI or Anthropic, according to one widely shared assessment.
I would genuinely love for this to happen
but many people think that OpenAI and Anthropic are already in a positive feedback loop
and as we have seen with Gemini 3 Pro: a ~5 trillion param reasoning model won't magically be AGI
(or for that matter a 6T param Grok-5)
my base case is that OpenAI and Anthropic will pull further ahead
xAI has less compute, less researchers, less data (no Codex, no Claude Code) and does not have access to models that literally speed up research (behind ~6 months)
Google on the other hand is still in the race, being only ~3 months behind. they have the most compute, researchers, an infinite money glitch and the data
Rollout execution matters too. OpenAI’s ability to expose new models across products rapidly is contrasted on X with slower or more restricted Gemini access:
A year or two ago everyone said Google had the compute to compete with OpenAI and Anthropic but now that they're actually getting close honestly I just don't see it
Gemini 3.1 Pro isn't even available yet for all Gemini CLI users meanwhile OpenAI takes like 1-2 days max to roll out new models across all their platforms
Also OpenAI lets us use our Codex usage limits in products like OpenClaw but Google considers that abusive
And I don't think Gemini CLI or Antigravity even have more users than Codex
Google needs more TPUs
Rapid rollout may not win a snapshot leaderboard, but it can win developers. A model unavailable through the required API, region, toolchain, or compliance boundary has no production value for that buyer.
Who fits OpenAI or xAI? OpenAI remains relevant for teams prioritizing multimodality, broad ecosystem integration, rapid distribution, or price-performance. xAI may fit latency-sensitive and real-time-information workloads. Their low market odds should not be read as low odds of commercial relevance.
Why do DeepSeek and Alibaba still attract volume near 0%?
Polymarket currently displays DeepSeek at 0% and Alibaba at 1%, despite approximately $125,201 and $124,578 in respective trading volume.[1] That is not necessarily contradictory. Traders can buy very cheap tail-risk positions, sell after price changes, or trade an outcome heavily before its price collapses.
Earlier in the cycle, DeepSeek was treated as a credible disruption story:
DeepSeek has set off panic in the AI world.
But OpenAI is still the king.
There's only a 17% chance DeepSeek will have the best AI model by Q2.
The distinction is between changing industry economics and finishing first on a specified leaderboard. An inexpensive open-weight model can force frontier vendors to reduce prices, improve distillation, or loosen deployment terms without ever taking the contract’s top ranking.
Open models also serve buyers that need self-hosting, data sovereignty, customization, or predictable marginal cost. Those are substantial markets, but they are not necessarily measured by a Western preference leaderboard or composite frontier index.[10]
A surprise win by DeepSeek or Alibaba would imply that markets had badly underestimated release velocity or misunderstood the resolution criteria. Traders currently assign that scenario near-zero probability—not literal impossibility.
Is “best AI model” in 2026 a useful category at all?
The contract needs one winner. The industry increasingly does not.
Practitioners describe a stack in which Claude handles coding, Gemini handles research and long context, GPT handles image or multimodal tasks, Grok addresses real-time or niche questions, and specialized providers handle voice.
I wouldn't say that - just now for what I need Claude wins. However:
I use @bot to manage personal stuff its more fun to talk too and handles things good enough
@GeminiApp is better at researching - not at turning researching into result but they have so much data they are good for fact checking things and research when writing
GPT Image 2.5 is best image model
Grok 4.7 is best at solving niche issues in fintech world from my experience and understanding compliance and bank documentation
@ElevenLabs wins in voice
So I am not biased to say @claudeai is best and that can't change. However I would argue currently right now in what I need of AI, which is mostly coding and marketing, Claude wins from my personal experience.
Another popular framework divides the market into execution modes rather than vendors:
The AI landscape in 2026 isn't a winner-take-all market.
It has fractured into 4 distinct execution modes:
1. Claude 3.7 Sonnet (Anthropic)
• Moat: Hybrid reasoning & codebase architecture
• Best for: Autonomous coding, complex refactoring, full-stack debugging
2. GPT-5 (OpenAI)
• Moat: Broad multimodal reasoning & ecosystem integrations
• Best for: Multimodal synthesis, enterprise API layers, general research
3. Grok (xAI)
• Moat: Real-time latency & uncensored data ingestion
• Best for: Breaking events, live sentiment tracking, unfiltered analysis
4. DeepSeek & Open Weights
• Moat: Zero API margin & self-hosted cost efficiency
• Best for: High-throughput local workflows, proprietary enterprise pipelines
Frontier benchmarks no longer matter. Choosing the right runtime for the specific task does.
That is closer to how production systems are built. Model quality varies by:
- programming and repository reasoning;
- factual recall and calibrated uncertainty;
- image, video, audio, and document understanding;
- context length and retrieval behavior;
- latency, throughput, and price;
- tool use and agent reliability;
- regional availability, privacy, and compliance.
Even the live benchmark summaries acknowledge that rankings vary by task.[8][9] A binary year-end contract flattens this multidimensional market into one resolvable result.
Data access further complicates the competition. OpenAI and Anthropic may benefit from visibility into user interactions, while Google’s privacy controls may restrict how teams inspect query data:
There is a simple reason why Gemini is so much worse than GPT or Claude
engineers at OpenAI or Ant can read incoming user queries. all the data is visible
but at Google there are tons of privacy restrictions preventing ppl from looking at data
basically building a model blind
That argument should not be treated as settled fact, but it identifies a real structural tension. User telemetry can accelerate product improvement; privacy restrictions can increase enterprise trust. The same constraint may be a disadvantage in training and an advantage in procurement.
What should founders, developers, and SaaS buyers do with these odds?
The actionable signal is not “choose Anthropic.” It is preserve optionality while the market remains contested.
Developers building AI features
Use an abstraction layer that can route tasks across providers. Maintain your own evaluation set based on actual user requests, including failure severity, latency, token consumption, and tool-call success. A 20-point Polymarket–Kalshi disagreement is a strong argument against hard-coding one provider throughout an application.
Early-stage founders
Start with the model that gets the product to market fastest, but separate prompts, evaluations, and business logic from provider-specific APIs where practical. If coding quality defines the product, Anthropic’s market leadership is relevant. If margins and response time define it, OpenAI, xAI, smaller models, or open weights may deserve more weight than leaderboard position.
Enterprise SaaS buyers
Evaluate the complete operating model:
- Capability: Does it succeed on your documents, workflows, and edge cases?
- Economics: What is the cost per successfully completed task—not merely per token?
- Deployment: Is the model available in the required regions and cloud?
- Governance: Can the vendor meet privacy, retention, audit, and access requirements?
- Switching cost: Can workloads move if rankings or prices change?
Google has a distinctive platform argument. If Gemini leads while being delivered through GCP and TPU infrastructure, the model advantage could improve the economics and strategic value of Google’s cloud stack:
Over the last few days, on Polymarket, $GOOGL's lead as the best AI model by the end of 2025 has increased even further.
Anthropic odds have also risen, while those of OpenAI and xAI have decreased.
While $GOOGL's market share in LLM search will be much lower than the market share it has on traditional search on the enterprise side if $GOOGL turns out to be the best model provider and on top of it offers them via GCP on their TPU infrastructure, GCP's value could be much more than the market current anticipates.
That is where year-end odds connect to SaaS industry structure. Model leadership can influence cloud consumption, developer tooling, marketplace distribution, and negotiated enterprise bundles. Yet the winning model contract will not reveal which provider offers the best total cost of ownership for a particular company.
How should practitioners read the smart money through December 31?
As of October 2, the market implies Anthropic leads at 54%, Google is the credible challenger at 36%, and OpenAI, xAI, DeepSeek, and Alibaba collectively occupy the long tail. The strongest market narrative favors Anthropic’s coding and agentic moat; the strongest counterargument favors Google’s long context, computer use, reliability, and infrastructure integration.
Before the expected December 31 resolution, practitioners should monitor:
- new Claude, Gemini, GPT, Grok, DeepSeek, or Qwen-family releases;
- movement on the leaderboard named in the contract rules;
- broad availability, not only benchmark announcements;
- changes in Polymarket odds after releases;
- whether the Polymarket–Kalshi gap narrows;
- cost, latency, and reliability on production-shaped evaluations.
The correct procurement conclusion is not that Anthropic will win. It is that traders currently consider Anthropic the most likely single leaderboard winner, while assigning Google a substantial chance of overtaking it. The market’s disagreement with other venues—and the industry’s fragmentation by task—means no prudent team should confuse that expectation with a durable architectural truth.
Use the odds to calibrate expectations. Do not use them to bet your entire stack on one lab.
Sources
[1] Polymarket — Which company has best AI model end of 2026?
[2] DeFi Rate — Best AI Model of 2026 Odds: Who Will Be #1 at Year-End?
[4] Casino.org — Best AI Model Odds: The Top Markets and Outcomes
[7] Polymarket — AI Predictions & Real-Time Odds
[8] Veso Research — Generative AI Model Ranking Matrix
[9] BenchLeader — LLM Leaderboard 2026
[10] BenchLM — LLM Leaderboard & AI Model Benchmarks, October 2026
[11] BenchLM — Frontier AI Models: Live Top 10 Rankings, Evidence and Pricing
[12] INDmoney — Gemini 4 Argon vs OpenAI & Anthropic: AI Race Explained
References (16 sources)
- Polymarket — Which company has best AI model end of 2026? - polymarket.com
- Best AI Model of 2026 Odds: Who Will Be #1 at Year-End? - defirate.com
- Polymarket Assigns 74 Percent Probability to Anthropic for Best AI Model at 2026 Close - sccgmanagement.com
- Best AI Model Odds - The Top Markets and Outcomes - casino.org
- Anthropic Wins Best AI Model: September 2026 Market Resolved - lines.com
- SK Hynix, Samsung rally lifts AI mood as Polymarket puts Anthropic at 98% - blockchain.news
- AI Predictions & Real-Time Odds | Polymarket - polymarket.com
- Generative AI Model Ranking Matrix · Veso Research - veso.ai
- LLM Leaderboard 2026: top AI models ranked | BenchLeader - benchleader.com
- LLM Leaderboard & AI Model Benchmarks — October 2026 - benchlm.ai
- Frontier AI Models: Live Top 10 Rankings, Evidence and Pricing (October 2026) | BenchLM.ai - benchlm.ai
- Gemini 4 Argon vs OpenAI & Anthropic: AI Race Explained - indmoney.com
- Announcing Artificial Analysis Intelligence Index v4.2 - artificialanalysis.ai
- LLM Leaderboard 2026 - Top AI Models Ranked | LM Market Cap - lmmarketcap.com
- The AI Rankings - theairankings.com
- 大模型排行榜 · LLM Leaderboard · 模型能力与成本对比 - raw.githubusercontent.com