The $3.6M Bet on Anthropic: What Polymarket's AI Model Odds Reveal for Developers in 2026
Polymarket AI model odds put Anthropic at 98% for best model in September 2026. Discover what $3.6M in bets reveals for developers, founders, and SaaS buyers.

The practical question for developers, founders, and SaaS buyers is not simply whether Anthropic will “win” September. It is whether Polymarket’s overwhelming preference for Anthropic should change model selection, product architecture, or vendor strategy.
The short answer: treat the odds as a strong forecast of one benchmark-based resolution—not as proof that Anthropic is the best supplier for every workload. As of September 19, 2026, traders price Anthropic at a 98% implied probability in a market with roughly $3,597,720 traded. That is meaningful evidence about the expected leaderboard winner around October 1, but it says much less about coding-agent reliability, inference economics, creative performance, or the long-term threat from cheaper open-weight models.[13]
Bottom line
>
- Anthropic: 98% — traders expect its current benchmark lead to survive through resolution.
- Google: 1% — the only alternative with a displayed probability above zero.
- OpenAI, Meta, SpaceXAI, and DeepSeek: approximately 0% each — the market currently sees little chance of a qualifying late upset.
- For practitioners, the result is a benchmark signal, not a procurement verdict. Test models on your workload, at the effort level and token budget you can actually afford.
What is the market actually pricing on September 19, 2026?
The Polymarket contract asks which company will have the best AI model at the end of September 2026 and is expected to resolve around October 1.[13] At the September 19 snapshot supplied for this analysis, traders price the field as follows:
| Company | Implied probability | Trading volume |
|---|---|---|
| **Anthropic** | **98%** | **$788,753** |
| **Google** | **1%** | **$312,946** |
| **OpenAI** | **0%** | **$671,102** |
| **Meta** | **0%** | **$402,339** |
| **SpaceXAI** | **0%** | **$345,254** |
| **DeepSeek** | **0%** | **$179,457** |
Total market volume is roughly $3,597,720. CryptoSlate and Lines.com also track the contract as a prediction market rather than a comprehensive evaluation of every model and deployment scenario.[14][15]
An implied probability is the price-derived market estimate of an outcome—not a scientifically measured likelihood. Displayed odds of 0% can also reflect rounding rather than literal impossibility. Prices move with available liquidity, trader positioning, new releases, and interpretations of the resolution rules.
Volume means something different. OpenAI’s $671,102 in traded volume shows substantial activity, but it does not mean $671,102 is currently betting for OpenAI. Every trade has two sides, and historical turnover can remain high even after the market has repriced an outcome sharply downward.
That distinction is central to reading the contract correctly:
Polymarket gives Anthropic a 91% shot at September’s “best” AI model. Sounds decisive—until the fine print: one text leaderboard settles it. Nearly $500,000 can price conviction. It can’t tell you which AI is best for your work. “Best” may be the laziest question in AI.
View on XThe market is making a narrow prediction: which company is most likely to satisfy the contract’s specified benchmark criterion on the resolution date. It is not determining which model will produce the best return on a SaaS company’s inference budget.
Why does the market imply a 98% chance for Anthropic?
The apparent basis for Anthropic’s lead is straightforward: Claude Fable 5.1 currently occupies the benchmark position most relevant to the contract. Anthropic’s September release record and model documentation provide the official context for its model lineup, while third-party leaderboards track comparative performance across providers.[1][2][6]
According to the live discussion around the launch, Fable 5.1 scored 66 on the referenced Intelligence Index, compared with 61 for OpenAI’s GPT-6 Astra under the same independent harness. That five-point gap is especially important because traders do not need to believe Anthropic is universally superior. They only need to expect that no qualifying competitor will displace it before the market resolves.
What the fff is this Anthropic & OpenAI document.
OpenAI put 99.9% on a slide and called it the start of the AGI era. ARC Prize re-ran the same benchmark on a standard harness and got 62.7%.
A neutral referee gave GPT‑6 Astra a 61 on the Intelligence Index, the same as its predecessor, after OpenAI’s biggest training run (100K+ GPUs), the score still didn’t budge.
Claude Fable 5.1 sits at 66 on the same harness, shipped two days earlier, same $10 in and $50 out, same 1M context.
> IMPORTANT: The whole document is a goldmine, seven rounds and the tests to settle it yourself in the article below
Astra's reasoning comes back encrypted and OpenAI's own system card says the model is harder to monitor than the last one. It also shipped rated Critical for cyber with 100% on ExploitBench.
Then the part nobody is saying out loud, the best Anthropic model is not for sale. Mythos 5.1, same weights as Fable with fewer guardrails, 60.9% on Terminal-Bench against Astra's 57.9%, handed only to verified labs.
Two flagships, 48 hours apart, everyone picked a side before reading a single benchmark, and the actual answer is that they are even.
This illustrates why standardized harnesses matter. A harness is the testing infrastructure that controls prompts, tools, scoring, retries, and model settings. Vendor-published numbers can differ from independent results when the test configuration changes. Traders therefore appear to be giving more weight to the leaderboard that governs resolution than to broader launch claims.
Anthropic also benefits from perceived continuity. The September 1 release of Fable 5.1 and Mythos 5.1 gave the market most of the month to absorb their performance, while benchmark trackers showed Anthropic in the leading position.[7][8] A challenger would need both a timely release and a score that qualifies under the contract’s rules.
The X conversation goes further, treating Anthropic as the target other labs now pursue:
I hate Anthropic more than anyone, but like it or not, their models are the industry standard.
Everyone used to chase Opus, and today they're chasing Fable.
Anthropic simply has the best data on earth.
Look at OpenAI: they flopped with the GPT-5 launch, fumbled around until 5.5 where things stabilized a bit, and then stumbled again with Sol, a reckless model that lacks human touch and real comprehension.
If you're a retail user paying $200 or less, OpenAI's models might be fine for you, but billion-dollar enterprises are all paying Anthropic.
There is no comparison.
That is an opinion, not independent evidence that every enterprise prefers Anthropic. But it captures the expectation embedded in the 98% price: traders currently see Anthropic as the incumbent benchmark leader, not the company attempting a last-minute catch-up.
Is the market underpricing OpenAI’s ability to catch up?
The strongest case against the consensus is not that OpenAI already leads the designated leaderboard. It is that OpenAI may have a faster iteration engine than a month-end snapshot can fully capture.
Some practitioners argue that Anthropic and OpenAI operate with broadly comparable compute resources, but that OpenAI extracts more performance through inference optimization:
anthropic has only ~20% less compute than openai and yes fable is definitely big but i think openai is just way, way better at optimizing inference for some reason
View on XOthers point to OpenAI’s rapid Codex cadence—5.1-Codex in November 2025, 5.2-Codex in December, and 5.3-Codex in February 2026—as evidence of a fast reinforcement-learning pipeline. Reinforcement learning, or RL, is the post-training process used to reward useful behavior and improve performance on targeted tasks.
First time seeing a representative of an AI Lab confirm that models are trained on their harness.
Doesn't mean it hasn't been mentioned.
But first seeing it for me.
Anthropic has been ahead with Claude Code because Claude Code came out of the gate first.
But OpenAI is catching up *FAST*.
My intuition is that OpenAI has the most rapid RL pipeline capability, which is why you saw such a rapid succession of:
> 5.1-Codex --> 11/12/25
> 5.2-Codex --> 12/18/25
> 5.3-Codex --> 2/5/26
If OpenAI hasn't already surpassed Anthropic and Opus 4.6 with GPT-5.3-Codex...
They certainly will with the next iteration.
The bull case is that OpenAI can continue “hill-climbing”: repeatedly training, evaluating, and optimizing against coding and agent tasks. If that pipeline produces a qualifying release before the cutoff, today’s near-zero displayed price could prove too pessimistic.
From April after the Mythos Preview and GPT-5.5 launches:
- "I think my scenario of OpenAI being ahead by 1-3 months by end of year is more likely after this launch"
- "it's quite likely now that OpenAI will out-accelerate Anthropic"
Anthropic obviously has a better model internally, but I don't think it's that much better and OpenAI just continued RL training for GPT-6.1-Astra
OpenAI is just starting to hill-climb, while Anthropic already had ~7 month of hillclimbing with a massive model
But there are three reasons traders may still discount that possibility.
First, a rapid development cadence does not guarantee a leaderboard-changing release within roughly 11 days. Second, improving coding performance does not necessarily lift the specific text-intelligence score that settles the market. Third, if labs train against their own evaluation harnesses, gains may transfer imperfectly to an independent harness.
There is also a product-level disagreement about specialization. OpenAI’s coding focus may create more useful agents for software teams, even if it does not produce the broadest model. Critics argue that optimizing too heavily for code can weaken writing, creativity, or strategic reasoning:
The gap between OpenAI and Anthropic's flagships is dead obvious.
Anthropic's models are sharper, broader, and actually think with you.
When OpenAI fell behind, they panicked and hyper-focused on coding just to stay relevant.
But they ignored a basic reality: train a model solely on code, and you kill its soul and creativity.
It becomes a sterile execution engine that any competitor can replicate.
Look at GLM-5.1, for instance. It was incredible, right on Opus's heels.
Then Zai's obsession with code at the expense of everything else completely trashed its cognitive range.
GLM-5.3 ended up useless outside programming, totally butchering writing, creativity, and strategic planning.
That is exactly what's happening to Sol right now, and what already happened internally with Astra.
On top of that, Anthropic's engineers are just far more competent, while OpenAI's devs got lazy and rely entirely on their own models to do the work.
By the way, what I'm saying about Astra comes directly from private hands-on testing and insider sources, not speculation. You'll all see soon enough.
For a developer choosing a coding agent, that trade could be entirely acceptable. For a marketing platform or research product, it might not be. The Polymarket price cannot resolve that distinction.
Is “best AI model” the wrong question for production teams?
For procurement, “best” is usually underspecified. The answer changes with the task, reasoning setting, latency target, context, tools, and acceptable error rate.
The effort-setting issue is particularly consequential. An independent test discussed on X reportedly scored Claude Fable 5.1 at 58 on Low effort and 66 on Max. Obtaining those eight points required 143.7 million output tokens instead of 13.1 million—about eleven times as many.
An independent lab ran Anthropic's top model through its full test at all five effort settings and published every score and every token count.
for free.
every AI newsletter you read told you to pick the smartest model and never mentioned the one setting that decides how hard it thinks.
the lab is Artificial Analysis. the model is Claude Fable 5.1, top of their Intelligence Index at launch.
on Low it scored 58. on Max it scored 66. to buy those 8 points it wrote 143.7 million output tokens instead of 13.1 million. eleven times the output.
the part nobody quotes is in Anthropic's own launch post: on Low or Medium, Fable 5.1 matches or beats Fable 5 at a much lower cost. then Claude Code ships on High by default.
OpenAI says it on its own help page for GPT-6 Astra: "Astra at Low effort can outperform Sol at High effort." and one line above it: "Lower effort does not mean lower capability across models."
most people paying for the top plan have never once opened the menu that decides how fast it runs out.
An effort setting controls how much inference-time computation a reasoning model is allowed to use. Higher effort can improve difficult answers, but it may increase token consumption, latency, and cost. A leaderboard that compares models at their maximum settings can answer “which configuration achieves the highest score?” while failing to answer “which configuration can economically serve 100,000 customer requests?”
Benchmark choice creates another distortion. Traditional coding tests often begin with a clear specification. Real agents may instead need to inspect an unfamiliar product, infer its behavior, identify missing workflows, implement them, and verify that nothing broke. ProgramDistill attempts to test that broader process:
New benchmark ProgramDistill: instead of giving coding agents a written spec, it hands them a working app and says "rebuild this."
→ 4,063 auto-generated tasks from 26 real web apps → Best agent (GPT-6 Astra) rebuilds only 49% of full workflows; Claude Opus 5 gets 29% → Success collapses from 100% → 64% as repair depth grows 1 → 8 → 59% of failures: the agent never even tried the feature it was supposed to copy
The bottleneck for AI coders isn't writing code. It's paying attention — and double-checking the right thing.
#agent #llm #benchmark
In the results summarized in that post, GPT-6 Astra rebuilt 49% of full workflows while Claude Opus 5 reached 29%. The more revealing detail is that 59% of failures reportedly occurred because the agent did not attempt the relevant feature. That suggests the limiting factor can be observation and verification—not code generation.
Leaderboard trust also becomes harder when labs optimize on familiar harnesses. A model can become excellent at a benchmark’s task distribution without producing equal gains elsewhere. Live benchmark collections help expose variation across intelligence, coding, and agentic tests, but no single score eliminates the need for workload-specific evaluation.[9][10]
The correct production question is therefore:
Which model, effort setting, toolchain, and context architecture meets our quality threshold at an acceptable cost and latency?
That question rarely has the same answer as a prediction contract.
What do the odds miss about cost, margins, and vendor lock-in?
Polymarket rewards the company expected to top the designated leaderboard. SaaS economics reward a different combination: sufficient quality, low cost-to-serve, predictable capacity, and durable access.
An X post citing Q2 2026 figures claimed that Anthropic booked approximately $11.6 billion with roughly $300 million in operating profit, while OpenAI booked $6.7 billion and recorded a $12.3 billion operating loss:
@OpenAI vs. @AnthropicAI: same playbook, very different numbers.
Q2 2026:
Anthropic: $11.6B booked, ~$300M operating profit
OpenAI: $6.7B booked, $12.3B operating loss
But one quarter doesn’t settle a multi-year enterprise buying decision.
Ray Rike & Peter Buchanan break down:
• Claude Code vs. Codex
• GTM leadership + enterprise continuity
• Channel conflict
• Distribution bets
• Hyperscaler competition
• Open-weight economics
The real battle: the same enterprise budget from multiple directions.
Those numbers should be treated as claims from the live industry conversation, not as independently established by the benchmark sources cited here. Their analytical significance is still clear: model leadership only becomes durable business leadership if the provider can serve demand economically.
Token allowances and overages are a particular concern for agentic software. Coding agents can read large repositories, invoke tools repeatedly, generate long traces, and retry failed operations. One practitioner estimates that heavy users could face $400 to $1,000 per person per day in overages:
Prediction: Claude has massively taken the lead right now because they offer a better product, but that comes at a massive cost.
Buyers have not realized that included in a Claude subscription is not enough tokens to get real work done and that overages will cost $400 to $1,000 per day per user. Anthropic will need to buy significantly more compute, but because they don't own their own data centers, the cost to serve will continue to go up.
Spend will shift gradually and then quickly back to OpenAI, who can offer comparable models but at a much lower cost basis because they own their own data centers. Cost of inference will become the only competitive advantage making this market a race to the bottom.
Apple or Google will buy or merge(!!!) with Anthropic.
That forecast is disputed and highly dependent on plan design and usage. Nevertheless, founders should model the scenario rather than assume a flat subscription represents the full cost of production-grade usage.
Access policy creates a second risk. Anthropic has reportedly restricted some third-party applications from using Claude subscriptions and cut access for rival labs:
Anthropic blocked Claude subs in third-party apps like OpenCode, and reportedly cut off xAI and OpenAI access.
Claude and Claude Code are great, but not 10x better yet. This will only push other labs to move faster on their coding models/agents.
DeepSeek V4 is rumored to drop soon, with stronger coding than Claude and GPT. Whale Code soon?
Competition is great for users.
Such restrictions can protect capacity, prevent account misuse, or preserve a product moat. For customers, however, they illustrate why access through a consumer subscription, first-party coding product, and metered API should not be treated as interchangeable.
The practical divergence is sharp:
- Anthropic may fit teams that value current benchmark leadership and can justify premium inference for high-value tasks.
- OpenAI may fit teams prioritizing rapid coding-agent iteration, broad distribution, or a potential cost advantage at scale.
- Google may fit organizations already committed to its cloud and productivity stack, even though traders assign it only a 1% probability in this particular contract.
- A multi-model architecture fits SaaS companies whose margins or continuity cannot tolerate one provider changing price, policy, or access.
Why do open-weight and Chinese models matter despite near-zero odds?
DeepSeek’s displayed 0% probability means traders currently see little chance that it will satisfy this contract by the end of September. It does not mean traders have disproved its longer-term economic threat.
The strongest challenge from DeepSeek and GLM is cost-adjusted capability. BridgeMind reports that DeepSeek V4.1 Flash beat GPT-6 Astra on one image task while costing $0.03 versus $0.59:
DeepSeek V4.1 Flash just beat GPT 6 Astra on the BridgeBench ocean sunset test. For 3 cents.
$0.03 vs $0.59. Twenty times cheaper. Faster too. And look at the two oceans. The DeepSeek one is better.
Five days ago I said OpenAI might kill Anthropic on cost.
Now a Chinese lab is doing to OpenAI what OpenAI did to Anthropic, at 1/20th the price.
DeepSeek might have actually cooked on this one.
One test cannot establish general superiority, and the “twenty times cheaper” comparison may not carry over to other workloads. But it demonstrates the pressure frontier providers face: customers do not need a cheaper model to win every benchmark. They need it to clear the quality threshold for a large share of routine requests.
A separate comparison of GLM-5.3-Flash and DeepSeek V4.1 Flash found split results across code review, mobile review, and planning. GLM won two of three tasks, while DeepSeek found more issues but also generated more false positives:
GLM-5.3-Flash vs DeepSeek-V4.1-Flash
I ran the exact same tasks on both on an agent harness I'm building:
- Code review
- Mobile review
- Improvement plan
Then I verified every claim from each model against the code. Who's the winner?
▶️ On code review, GLM 5.3 Flash won.
GLM finished in 27 minutes and worked by reading 35 file reads, and performed 27 searches.
DeepSeek took 45 minuted and worked by running 55 shell commands.
DeepSeek raised 23 issues, GLM raised 10.
Both were right about roughly 60% of the findings. So DeepSeek found more real problems, and more fake ones.
That's not what decided it though.
GLM's biggest finding was correct. DeepSeek's biggest finding was wrong, and that is why GLM 5.3 Flash won on code review.
▶️ On mobile review, DS4.1F won.
It found a bug that makes the composer unusable the second messages start queuing. GLM had that exact component on screen, asked the wrong question about it, and moved on.
That said, my prompt explicitly instructed both models that look and feel were super important. GLM viewed 14 screenshots. DeepSeek viewed zero.
▶️ What about the improvement plan? This one goes to GLM 5.3 Flash.
GLM was broader and better prioritized. It also came back with 33 improvements, each one sized and properly prioritized.
DeepSeek v4.1 Flash came back with 14. Deeper on a couple of them, but it spent almost nothing on UI and performance. Two of the four things I asked for.
On this task, both were accurate. I verified every claim.
So which is better?
🥇 GLM 5.3 Flash takes it 2-1. Better calibrated, roughly 1.3-1.7x faster, and it knows when its own tooling is the problem.
🥈DeepSeek v4.1 Flash feels "sharper" and the one most likely to be confidently wrong about what it just measured.
For me GLM 5.3 Flash is the one I run daily.
That is more representative of real procurement than a universal ranking. A security team may prefer higher recall and tolerate manual review. A small engineering team may favor better calibration because false positives consume scarce developer time.
The broader thesis is that open-weight models may trail frontier systems by only months, not years:
Wat? The open source / open weight AI models are only 4-6 months behind the frontier models in capability.
The more capable and inexpensive they become, the less that Anthropic / OpenAI / https://X.ai will be able to charge.
How will they sustain trillion dollar valuations if their competitors charge a fraction of their prices for a product that is as good or better?
If they don't erect a regulatory moat for themselves now, the open weight / open source inference providers will eat their lunch.
If that expectation holds, frontier pricing power could weaken even while Anthropic continues to top selected leaderboards. Open-weight deployment can also offer data control, customization, and predictable infrastructure costs, although it transfers operational responsibility to the customer.
This is where the September market has the least to say. A one-month contract heavily discounts developments that affect 2027 margins rather than an October 1 resolution. DeepSeek can be correctly priced near zero for the contract and still be highly relevant to a SaaS company’s architecture.
What should developers, founders, and SaaS buyers do with these odds?
The 98% Anthropic price is useful as a directional signal of current benchmark confidence. It is not a reason to rewrite a production stack without further testing.
Developers: build a small internal evaluation harness
Test representative tasks using the same prompts, tools, context, and effort settings expected in production. Record:
- Task completion and correctness
- False-positive and hallucination rates
- Tokens and cost per successful task
- End-to-end latency
- Tool-call failures and retries
- Performance at Low, Medium, High, and Max effort where available
Do not test only isolated code generation. Include repository navigation, requirement discovery, verification, and repair. Current benchmark collections can help identify candidates, but internal tasks should make the final decision.[8][9]
Specialized context may also reduce dependence on the nominally strongest base model. Google’s Chrome team, for example, reportedly improved best-practice adherence by supplying current web-platform guidance through an MCP server:
Google's Chrome team published guidance for 128 modern web platform features. Apache 2.0. Installs as an MCP server into Claude Code, Gemini CLI, Copilot CLI, Goose, and Vercel.
The benchmark result they published: a 37 percentage point improvement in best-practice adherence on web platform tasks. 10,000 installs since launch.
The mechanism is interesting. Coding agents hallucinate on web platform APIs because their training data under-represents recent specifications. A guidance layer installed as an MCP gives the agent current, curated patterns for each feature without retraining. The agent calls the MCP, gets the right patterns, and uses them.
This is the same pattern as Greptile's Knowledge Base MCP, but for web platform rather than codebase structure. Specialized context delivered at inference time instead of baked into training.
The more MCP servers like this ship, the less model quality matters on domain-specific tasks. The bottleneck shifts from knowledge to retrieval architecture.
For domain-specific systems, retrieval architecture and curated context can matter more than a small difference in base-model scores.
Founders: avoid single-provider assumptions
Use an abstraction layer that can route requests by task, quality requirement, and price. Preserve access to prompts, traces, evaluation data, and tool schemas so migration is possible.
A sensible pattern in late 2026 is:
- Route the hardest reasoning tasks to the current frontier leader.
- Send routine extraction, classification, and drafting to cheaper models.
- Maintain a tested fallback for outages or policy changes.
- Re-run evaluations after every major model or pricing update.
This does not require supporting every provider. Two genuinely interchangeable paths are more valuable than six nominal integrations that have never been tested under failure.
SaaS buyers: evaluate total workflow cost, not token price alone
A cheaper model can be more expensive if it creates retries, bad recommendations, or human review. Conversely, a premium model can destroy gross margin if high-effort reasoning is used on every request.
Calculate cost per accepted outcome, including inference, retries, tooling, review, and latency. Negotiate around rate limits, data retention, model deprecation, and third-party access—not only headline token prices.
The final interpretation is therefore deliberately narrow: betting markets currently put Anthropic’s odds at 98% for this September 2026 resolution, and that strongly implies traders expect its leaderboard lead to hold. Developers and buyers should use that signal to decide what to evaluate first, not what to adopt automatically. The AI market is moving toward fragmentation: premium frontier models for the hardest tasks, cheaper open or specialized models for volume, and orchestration layers that make “best model” less important than “best system.”
Sources
[1] Anthropic — Model system cards
[2] Claude Help Center — Release notes
[6] Releases Index — Anthropic Release Notes & Changelog, September 2026
[7] BenchLM.ai — LLM Leaderboard & AI Model Benchmarks, September 2026
[8] BenchLM.ai — Who Is Winning the AI Race? Monthly LLM Leader Timeline
[9] MultipleChat — Live AI Benchmarks
[10] LMSpeed — Model Performance Leaderboard
[13] Polymarket — Which company has the best AI model end of September?
[14] CryptoSlate — Which company has the best AI model end of September odds and analysis
[15] Lines.com — Which Company Has the Best AI Model in September 2026?
References (15 sources)
- Model system cards | Anthropic - anthropic.com
- Release notes | Claude Help Center - support.claude.com
- Claude (AI) - Wikipedia - en.wikipedia.org
- Anthropic Claude | endoflife.date - endoflife.date
- Claude Models & API IDs: Current Anthropic Model List | BenchLM.ai - benchlm.ai
- Anthropic Release Notes & Changelog · September 2026 — Releases Index - releases.sh
- LLM Leaderboard & AI Model Benchmarks — September 2026 - benchlm.ai
- Who Is Winning the AI Race? Monthly LLM Leader Timeline (September 2026) - benchlm.ai
- Live AI Benchmarks — Intelligence, Coding & Agentic Leaderboards — MultipleChat - multiple.chat
- Model Performance Leaderboard | LMSpeed - lmspeed.net
- Best AI Models in 2026: The Complete Ranking - theairankings.com
- AI Benchmarks 2026 - MMLU, GPQA, SWE-bench | LM Market Cap - lmmarketcap.com
- Which company has the best AI model end of September? Trading Odds & Predictions 2026 | Polymarket - polymarket.com
- Which company has the best AI model end of September Odds & Prediction Market Analysis | CryptoSlate - cryptoslate.com
- Which Company Has the Best AI Model in September 2026? Winner Odds | Lines.com - lines.com