Epoch AI: 2027 Chip Shipments Could Support 30–170M Frontier Model Agents
Key point
Epoch AI estimates that HBM hardware shipped through 2027 could support 30–170 million simultaneous frontier model agents, equivalent to the weekly work of 140–720 million full-time employees.
Details
Epoch AI estimates that high-bandwidth memory (HBM) hardware shipped between 2025 and 2027 could support 30 million to 170 million simultaneous frontier model agents if fully deployed and allocated to agent workloads. This capacity translates to the weekly output of approximately 140 million to 720 million full-time employees, assuming agents run 168 hours per week.
Capacity Estimates and Hardware Constraints
The analysis uses HBM as the primary supply constraint, noting that memory capacity limits concurrent requests while bandwidth limits streaming speed. Key findings include:
- 2025–2026 Shipments: Estimated to support 16 million to 56 million simultaneous agents.
- 2027 Shipments: Cumulative capacity rises to 30 million to 170 million simultaneous agents.
- Efficiency Gains: Applying benchmarks for DeepSeek V4 Pro suggests capacity could reach 1.9 billion simultaneous agents, equivalent to 8 billion weekly full-time workloads.
The study assumes HBM4/4E systems double the concurrent agent capacity per GB compared to HBM3E systems, though sensitivity analysis considers 1x and 4x improvement scenarios. Major GPU architectures like Nvidia Blackwell Ultra (GB300) and Rubin (VR200) are modeled with 288 GB of HBM, with Rubin offering 2.75x the bandwidth of Blackwell.
Economic Implications and Supply-Demand Mismatch
A central concern is the potential for a massive supply-demand mismatch. Even if only 20% of central capacity estimates are utilized, the implied annual API-equivalent spending would range from $2.6 trillion to $5.3 trillion. This figure significantly exceeds projected model developer revenues, which are estimated at approximately $1 trillion by the end of 2027 assuming a 5x annual growth rate.
The report highlights several factors influencing this gap:
- Deployment Delays: Physical infrastructure constraints, such as power and data center space, may delay chip installation, allowing demand to catch up.
- Cost Decline: The cost to achieve specific AI performance levels has dropped by 47% per quarter since 2023, potentially lowering the revenue required to justify hardware investments.
- Jevons Paradox: Increased efficiency may drive higher total demand for inference compute as users delegate more complex tasks to persistent agents.
Methodology and Data Sources
Epoch AI derived these estimates using two primary data sources:
- Open Models: Benchmarks from SemiAnalysis’s AgentX, measuring concurrent sessions at P90 streaming speeds (50–100 tokens per second).
- Closed Models: Analysis of TraceLab public datasets (Codex, Claude Code logs) to estimate API-equivalent costs per agent-hour, ranging from $15.50 for GPT-5.6 Sol to $50.20 for Fable 5.
The study assumes a baseline serving cost of $5.00 per GB300-hour and a revenue-to-cost ratio of 5x to 10x, resulting in estimated serving costs of $3.00 to $6.00 per agent-hour.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.