The price floor dropped. Mid-quarter.
On September 22, 2026, OpenAI launched GPT-6 Sol and Luna with API prices roughly 50% below GPT-5.6’s promotional rates. The same day, Anthropic released Claude Opus 5.5 at approximately 40% lower cost per typical workload than Opus 5. Then on October 4, at DevDay 2026, OpenAI announced GPT-6.1 Sol at near-Astra-level intelligence for one-fifth of Astra’s standard token prices. Source: OpenAI DevDay 2026 Recap. Three repricing events inside six weeks. That is not a promotional cycle — it is a structural shift in what frontier AI costs.
Box CEO Aaron Levie called September 22 “an insane day” and noted that “the rate at which cost per task drops in AI is unlike any other type.” Source: Yahoo Finance, Sep 24 2026. The more precise observation for you, as a founder or executive: the benchmark winner now costs a fraction of what last quarter’s budget assumed. Enterprises still anchored to 2025 procurement baselines are systematically overpaying.
What the numbers show
GPT-6 Sol launched at $2/$10 per million input/output tokens. GPT-6 Luna launched at $0.10/$0.50 per million tokens. On AutomationBench 1.0.6, GPT-6 Sol scored 33.2% at $0.27 per task. Astra scored 30% at four times that cost. Opus 5 scored 27% at eleven times the cost. The top benchmark performer is now also the most cost-efficient at scale — a combination that did not exist a quarter ago. Source: AI Socratic, Sep 21 2026.
On Agents’ Last Exam, GPT-6 Sol scored 56.4%, per AI Socratic’s September 2026 data. Source: AI Socratic, Sep 21 2026. On DeepSWE 1.1, Sol scored 68.8% — just 1.1 points behind Fable 5 — at one-fifth of Fable 5’s price. Luna scored 66.6% on the same benchmark. These are not marginal efficiency gains. They represent a materially different cost-per-outcome curve.
Claude Opus 5.5 entered at $4/$20 per million input/output tokens, down from $5/$25 for Opus 5. Cache-read pricing dropped 60% to $0.20 per million tokens. Anthropic claims Opus 5.5 outperforms GPT-5.6 Sol on a software-development benchmark at roughly one-third the per-task cost. Full independent verification had not been published as of September 28, 2026. Source: shattered.io, Sep 28 2026. Treat that claim as a strong signal, not a settled fact, until third-party benchmarks confirm it.
Fable 5.1 cut cache-read pricing from $1.00 to $0.25 per million tokens — a 75% reduction yielding effective 25–45% total savings on cached workloads. Source: Local AI Zone, Sep 3 2026. Across the top-15 models as of September 2026, the price spread runs from Fable 5.1 at $11.90 blended to Muse Spark 1.3 at $0.10 — a 119× range. Model selection is now a primary cost-engineering decision, not a capability preference.
Why this happened now
The proximate cause is competition for production-scale enterprise contracts. Deloitte’s 2026 survey reports that 67% of enterprises have moved at least one AI agent initiative from pilot to production, up from 23% in 2024. Source: Deloitte State of AI in the Enterprise 2026. When buyers commit to volume at scale, margin per token matters less than locking in long-term contracts. Providers are competing on cost-per-task, not capability headlines.
The structural cause is inference efficiency. Each model generation requires less compute per output token than the last. Providers are passing a portion of those gains to buyers — mostly through competitive necessity. Cost-per-task is falling faster than most enterprise software budgets anticipated.
Who this affects and where the limits are
This repricing matters most to enterprises running high-volume agentic workloads where token costs compound. McKinsey data, as reported by MarketScale, indicates 60% of agentic AI costs go to response refinement alone. Source: MarketScale citing McKinsey. A 40–60% reduction on the dominant cost line changes unit economics materially. It also matters to procurement teams who locked in annual AI spend before September 22, and to founders building AI-native products where model cost is a direct input to gross margin.
But the limits deserve equal attention. Lower list prices do not automatically reduce total spend. Deloitte’s January 2026 report found that 74% of organisations hope to grow revenue through AI in the future, versus only 20% already doing so. Source: Deloitte State of AI in the Enterprise 2026, Jan 20 2026. Lower cost-per-task typically expands consumption. If governance and budget controls do not account for volume growth, unit savings can disappear into higher aggregate spend. A lower capture rate on those savings is the realistic outcome for organisations without usage controls in place — not a stronger procurement argument.
Three actions before year-end
First, re-run your cost-per-task model using September 2026 pricing for GPT-6 Sol, Luna and Opus 5.5 against actual production workload profiles. The 119× spread across top-15 models means the model you defaulted to in 2025 is unlikely to be the right choice today.
Second, separate cached from uncached token spend before applying new rates. The cache-read cuts on Opus 5.5 (−60% to $0.20/M tokens) and Fable 5.1 (−75% to $0.25/M tokens) are disproportionately valuable for workloads with repetitive context — legal review, code review, document processing. Blending these into a single per-token average will understate the savings available to those workload types.
Third, pressure-test any annual AI contract signed before September 22 against current list prices. The repricing landed mid-contract-cycle for many organisations. Model stability, integration cost and vendor reliability all factor into a rational renegotiation — but the cost gap is now wide enough to warrant the conversation before year-end rather than at renewal.
Join the conversation or subscribe for weekly breakdowns.
Eagentix helps growth-focused enterprises redesign and automate manual business processes. We combine executive strategy, implementation support, and managed services to build dependable operations across Southeast Asia.
Eagentix helps growth-focused enterprises redesign and automate manual business processes. We combine executive strategy, implementation support, and managed services to build dependable operations across Southeast Asia.
Sources
- – The State of AI in the Enterprise – 2026 AI report | Deloitte US
- – September 2026 AI Model Updates: Every Launch, Price Move, and Architecture Shift – Local AI Zone
- – AI Socratic September 2026 — Jev AI Classifiers, Astra, and Pace the Frontier | AI Socratic
- – Claude Opus 5.5 Beats GPT-5.6 Sol at Third the Cost
- – DevDay 2026 Recap
- – Preparing for the Age of AI – A Living Outlook for Decision-Makers
- – OpenAI’s GPT-6 Just Got 50% Cheaper — Box CEO Says AI Agent Opportunity Could ‘Dramatically’ Expand: ‘What an Insane Day’
- – Smothering Heights
- – Gemini 4 Argon vs GPT-6 Astra: Benchmarks, Pricing
- – OpenAI GPT-6, Anthropic Opus 5.5 cut AI prices 50% | Value Add Pulse
- – 60+ AI Coding Model Stats for 2026 (Updated July 2026)
- – 60% of agentic AI costs go to response refinement: McKinsey
- – Which AI model to use? (Sept 2026) — IT Pro Expert
- – AI Usage Statistics 2026: Who Uses AI and How Much
- – GPT-6 Astra Benchmarks: Is It Better Than GPT-5.6?
- – GPT-6 Astra vs Gemini vs Claude: AI Agent Comparison
- – The Cost Floor Is Dropping: What GPT-6 Sol, Luna, and…
- – Enterprise AI Adoption in 2026: The Framework That Scales | Tommaso Maria Ricci
- – Gemini 4 Argon vs Claude Opus 5.5 vs GPT-6.1 Sol: 12-Pt Gap [2026] – Tech Insider Ireland
- – The Cost Curve
- – The 2026 Global Intelligence Crisis – Citadel Securities
- – AI Statistics & Trends 2026: Market, Adoption & Growth Data
