The margin lives in the infrastructure, not the model
Goldman Sachs projects token consumption will multiply 24 times by 2030, reaching 120 quadrillion tokens per month, driven by consumer and enterprise AI agent adoption. Goldman Sachs Research, May 2026. That number is not a forecast about model providers. It is a forecast about infrastructure — and infrastructure is where the margin lives.
The non-obvious implication: open-weight models are not disrupting hyperscaler pricing power. They are reinforcing it. As model-provider margins compress, infrastructure margins hold. Every token from every open-weight model still runs on hyperscaler GPUs, inside hyperscaler data centers, behind hyperscaler managed APIs. The margin shifts up the stack rather than disappearing.
What the margin data shows
AWS has historically operated at roughly 35–38% operating margins on infrastructure. Uncover Alpha, June 2026. Google Cloud posted operating margins above 33% in Q1 2026, and those margins continue to climb. Uncover Alpha, June 2026. Both figures hold regardless of which model generates the token.
Open-weight models — GLM 5.2, DeepSeek V3.2, Qwen3 Coder, Kimi K2.5, Llama, MiniMax M2.5 — are eroding model-provider brand premiums toward near zero. Uncover Alpha, June 2026. But the infrastructure bill is model-agnostic. As Jim Schneider noted in Goldman Sachs Insights: “The margin inflection for hyperscalers and model providers is very different from the prevailing market narrative that AI usage will simply drive an increasing and unsustainable cost burden.” Goldman Sachs Insights. The cost burden lands on the enterprise. The margin lands on the hyperscaler.
Why agents amplify consumption rather than reduce it
A chatbot query consumed a few thousand tokens. A single coding-agent session now chews through millions of tokens of context. Uncover Alpha, June 2026. The mechanism is demand elasticity: cheap inference removes the incentive to ration. You run the agent in a loop. You let it read the whole codebase. You re-run it five times and vote on the answer. Each iteration is a billable event.
So cheaper-per-token does not mean cheaper overall — it means more tokens consumed at a lower unit price, with total spend rising. Citadel Securities’ global macro analysis found that incremental compute capacity is finding paying demand rapidly, with every hyperscaler reporting that demand exceeds supply as of early-to-mid 2026. Citadel Securities.
The energy constraint reinforces this. The U.S. Department of Energy projects AI data centers could account for up to 12% of U.S. electricity consumption by 2028. Percepture, June 2026. Power supply constraints keep new entrants out and existing capacity scarce, sustaining infrastructure pricing power even as model costs fall.
The orchestration layer is the real lock-in
Model switching is now routine. Enterprises swap foundation models the way they swap database engines — with effort, but without rebuilding the stack. What they do not swap easily is the orchestration layer: the RAG pipelines, agent scaffolding, eval suites, and compliance controls built inside the cloud environment. Uncover Alpha, June 2026.
The reason is practical, not contractual. Data, security perimeters, billing, and compliance already live in the cloud environment. Building the agent harness there is the path of least resistance. Uncover Alpha, June 2026. But each integration deepens the switching cost. The hyperscaler captures orchestration margin on top of infrastructure margin — two layers of pricing power from a single architectural decision.
This is why Google is putting AI agents at the heart of its enterprise money-making push. Reuters, April 2026. Agents are not a product feature. They are a consumption-growth mechanism embedded in the layer enterprises are least likely to migrate away from.
Where the economics break unevenly
Not every agentic use case pencils out. Goldman Sachs analysis identified at least one real-time voice agent scenario where human labor cost was lower than the LLM cost, due to latency characteristics and time dependency. Goldman Sachs, May 2026. Agentic economics are highly task-specific, and a lower return on a given use case does not become a stronger case for deployment — it is a genuine constraint that narrows which tasks justify the spend.
AI monetization is broadly shifting from flat subscription models toward outcome-based and usage-based pricing. Seeking Alpha, May 2026. For the hyperscaler, how the SaaS layer prices its customers is irrelevant — the infrastructure bill is consumption-based regardless. The risk of misaligned pricing lands on the builder, not the platform.
Sovereign AI platforms argue that data residency without stack control is incomplete sovereignty. Lyzr. That framing has merit at the architectural level. But for most enterprises in 2026, the compliance and integration advantages of staying inside a hyperscaler environment are concrete, while the sovereignty benefit remains theoretical for their specific risk profile. The lock-in is chosen, repeatedly, for legitimate operational reasons — which makes it durable rather than fragile.
Implications for operators building on hyperscaler platforms
Treat the orchestration layer as a strategic asset, not plumbing. The RAG pipeline and eval suite you build today are the switching cost your vendor is pricing against tomorrow. Abstraction layers between your agent logic and provider-specific APIs preserve future optionality. Building directly against proprietary orchestration APIs is a decision with a long tail of consequences.
Model token consumption per agent task before committing to a use case. A task with high loop counts and large context windows will consume tokens at a rate that surprises teams accustomed to chatbot-era cost models. CloudZero and BCG’s 2026 analysis both provide frameworks for pre-deployment cost modeling. Run those numbers before you scale, not after.
Budget for total token volume growth, not unit price decline. Cheaper inference typically increases consumption. The Goldman Sachs 24x projection is a system-level forecast — your workload will reflect some version of that elasticity at the task level. A lower capture rate on a given use case lowers return; falling model prices do not automatically offset that.
Consumption pricing persists because the underlying economics reward it at the infrastructure layer, and because enterprises keep choosing to deepen their orchestration investment inside that layer. The conditions that would change this — portable orchestration standards, credible multi-cloud agent runtimes — are not imminent. Plan accordingly.
Eagentix helps growth-focused enterprises redesign and automate manual business processes. We combine executive strategy, implementation support, and managed services to build dependable operations across Southeast Asia.
Sources
- – Why Token Optimization Is a Gift to the Hyperscalers
- – AI Agents Forecast to Boost Tech Cash Flow as Usage Soars | Goldman Sachs
- – Agents Over Bubbles – Stratechery by Ben Thompson
- – Hyperscaler Kings: AI Data Center Leaders | Percepture
- – White House eyes data center agreements amid energy price spikes – POLITICO
- – AI monetization shifts to usage-based pricing (HUBS:NYSE) | Seeking Alpha
- – Top 5 Sovereign AI Platforms: A Buyer’s Comparison
- – Elastic Expectations
- – Pricing Strategies: Complete B2B Growth Guide
- – Google puts AI agents at heart of its enterprise money-making push | Reuters
- – Where You Should Build Your AI Agents – Smart Customer Service
- – AI Agent Cost: What Agents Really Cost To Run
- – AI Agent Platforms Compared: From Enterprise to Self-Hosted – Cuttlesoft, Custom Software Developers
- – Cost Management for AI Applications: Predictable Pricing vs. Usage-Based Billing
- – Why the AI capex cycle may just be beginning | CoBank | Cooperative. Connected. Committed.
- – As Seat-Based Pricing Fades, AI Startups Find Alternatives | ThinQ by EQT
- – The Top 10 AI Agent Platforms for Enterprise: A Comprehensive Comparison Guide
- – AI Agent Management Platform: Architect’s Guide
- – The commoditization of AI models and compute
- – Cloud AI Costs: Beyond Token Pricing Decisions | BCG
- – Who pays for AI’s electricity? Data centers spark debate over rising power costs
- – Consumer Agents Signal New Phase for AI Growth | Goldman Sachs
- – AI Agent Platforms in 2026: Comparison & Buyer’s Guide | Quickchat AI – AI Agents
- – AI Agent Development Platforms Compared (2026 Enterprise Guide) | assistents.ai
- – Why SaaS Companies Must Transform Their Business Models for the Agentic AI Era
