The CPU Comeback Is Upon Us

Amazon Web Services leaders reportedly instructed engineers earlier this year to conserve CPU cycles at all costs, as the company experienced an explosion in wait times for CPU server capacity while AI workloads strained its cloud infrastructure, IEEE Spectrum reports.
The AI boom initially drove demand for GPUs and later memory, while CPUs were largely excluded because their relative lack of parallelization made them a poor fit for AI model inference. The rise of agentic AI systems, which let models operate autonomously and call on sub-agents, is changing that, according to the report.
Matt Kimball, vice president and principal data center analyst at Moor Insights & Strategy in Austin, Texas, said 2026 has brought a spike in CPU demand, much of it due to agentic AI, describing enterprises scaling from 100 agents to tens or hundreds of thousands.
Intel senior staff research scientist Souvik Kundu said CPUs handle parsing output, determining which tool to invoke, making API calls or running code, collecting results, and feeding them back. AMD vice president Madhu Rangarajan said that in AMD's testing, seven of eight stages in realistic agentic AI pipelines run entirely on the CPU.
Kundu co-authored a paper with Georgia Tech researchers finding the CPU is often idle during GPU inference and the GPU often idle during CPU tool calls. They propose scheduling optimizations that can cut end-to-end latency by up to 1.8-times under sustained load.
Georgia Tech PhD student Euijun Chung co-authored another paper finding that too few CPU cores causes servers to fall behind dispatching work to GPUs, stalling GPUs. The paper covers tokenization, which Chung said must reprocess the whole sequence at every agentic tool call. It found time-to-first-token latency can rise dramatically with sequence length, and that increasing CPU core counts can reduce it by roughly 1.5 to 7 times at longer sequence lengths. Testing was limited to smaller models including Alibaba's Qwen 3-30B and Meta's Llama 3.1-70B.
Intel has sold out of server CPUs through at least the end of the year, AMD has doubled its server CPU forecast, Arm and Qualcomm have announced new CPUs for agentic AI, and Nvidia has prioritized Vera, its Arm-based CPU for agentic AI in the Vera Rubin platform. Kimball said this may lead to broader CPU shortages and higher prices, and noted Intel cut client CPU production in favor of server CPUs.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Earlier this year, leaders at Amazon Web Services delivered a new mandate to their engineers: They need to conserve CPU cycles at all costs. AWS has reportedly experienced an explosion in wait times for CPU server capacity as AI workloads strain the company’s cloud infrastructure. The issue seemingly took AWS off guard, and for good reason. The AI boom led to a surge in demand for GPUs and, later, memory . CPUs were mostly left out of the story, as their relative lack of parallelization made them a poor fit for AI