Microsoft and Amazon Web Services’ latest earnings reports send the same message: Arm-based CPUs have become the core infrastructure for agentic AI.

Recent earnings conference calls from Microsoft and Amazon jointly confirm a trend: custom chips based on the Arm architecture have shifted from earlier technical experiments to becoming a core pillar of hyperscale cloud infrastructure.

Microsoft CEO Satya Nadella summarized this architectural shift in one sentence: “When running agent workloads, the CPU is just as important as the GPU.” He then provided strong evidence: Microsoft expects that by the end of July, racks equipped with Azure Cobalt 200 will be deployed across more than 25 data centers worldwide.

A day later, Amazon announced that its chip business—including Graviton, Trainium, and Nitro—has surpassed an annualized revenue run rate of $25 billion, with year-over-year growth reaching triple digits. Customer purchase commitments for Graviton have nearly tripled quarter over quarter, and adoption of Graviton5 is also nearly twice as fast as Graviton4.

Although the products differ and the cloud platforms differ, the direction is converging: the world’s leading hyperscale cloud providers are turning Arm-based CPUs into the core of AI infrastructure and optimizing them for their own workload requirements and economic goals.

Agents need more than just model compute

The first wave of generative AI was measured by model size, the supply of AI accelerators, and token generation capacity. Agentic AI, by contrast, operates in a much more complex way: agents need to retrieve data, call tools, run code, execute policies, and verify results, involving a much broader range of tasks.

GPUs still handle the vast majority of model training and inference token generation, while CPUs handle the tasks surrounding model execution, including request routing, data preprocessing, tool orchestration, memory and storage management, and ensuring that thousands of concurrent interactive tasks keep running stably.

Nadella said on the earnings call that the virtual machine (VM) portfolio based on Azure Cobalt is currently supporting Microsoft’s own services as well as customer workloads from Adobe, Arm, Elastic, OpenAI, Sprinklr, and TomTom, showing that it has already been applied in real production environments.

Amazon Web Services is also confirming the same trend from another angle. Graviton5 is designed for agentic AI workloads and is optimized for scenarios such as real-time inference, code generation, and multi-step orchestration. Compared with the M8g instance based on Graviton4, the M9g instance based on Graviton5 delivers up to 25% higher compute performance and up to 35% faster machine learning inference.

Two hyperscalers, different paths to scale

The success of Cobalt and Graviton proves that Arm architecture enables hyperscale cloud providers to customize chips for their own massive infrastructure requirements while sharing the same architecture and software. This helps accelerate commercialization and generate business returns. Microsoft released Cobalt 200 in November 2025 and opened preview access to the VM product in June 2026. In less than two months, rack solutions equipped with the chip had been deployed in more than 25 data centers worldwide, an exceptionally fast pace.

Built on Arm Neoverse CSS V3, Cobalt 200 is Microsoft’s second-generation Arm-based cloud CPU. Compared with the previous generation, performance is up to 50% higher, it scales to 128 vCPUs, and it is designed for cloud-native, data-intensive, and agentic workloads.

Amazon Web Services has launched the fifth generation of Graviton products. M9g and M9gd instances powered by Graviton5 were officially launched into commercial use in June this year, with compute performance up to 25% higher than Graviton4. AWS says that across the Graviton product line, its price-performance ratio is 30% to 40% better than comparable instances.

What is truly worth noting is customer adoption of Graviton: 98% of Amazon Web Services’ top 1,000 EC2 customers are using Graviton, and more than 120,000 customers have built services on Graviton. For the third consecutive year, more than half of AWS’s new CPU capacity has used Graviton.

Efficiency unlocks compute, and compute creates revenue

Microsoft directly links infrastructure efficiency with business growth. Azure revenue rose 43% year over year in the quarter, while market demand still exceeded existing supply. Microsoft CFO Amy Hood said that continued improvements in CPU and GPU infrastructure operating efficiency, along with faster infrastructure expansion and activation, unlocked more usable compute resources, and those new compute resources were absorbed by market demand and converted into revenue almost as soon as they came online.

For hyperscale data centers, improving CPU efficiency means far more than just lower power bills. When market demand exceeds supply, more efficient infrastructure can unlock more usable compute at the same rack footprint and power budget, and that compute ultimately turns into revenue.

Microsoft did not disclose Cobalt 200’s specific contribution to efficiency gains, but the logic behind it is already very clear: performance, resource utilization, and energy efficiency are both levers for unlocking more compute and levers for driving business growth.

AWS provided another supporting data point. The $25 billion annualized revenue figure covers the entire chip business, not just Graviton. However, the rapid growth in purchase commitments and the accelerated adoption of Graviton5 suggest that the custom CPU line is becoming an important pillar of this large and fast-growing chip business. At the same time, its significant price-performance advantage provides customers with a strong reason to migrate more workloads to the Graviton platform.

Industry pressures continue to intensify. International Energy Agency data shows that AI-oriented data center power consumption grew 50% in 2025. Research firm Dell’Oro Group reported that in the first quarter of 2026, data center capital expenditures by Amazon, Google, Meta, and Microsoft rose 78% year over year. Compared with a year ago, every watt of power and every rack now matters more. CPU energy efficiency is no longer a secondary specification; it is a key means for hyperscale cloud providers to unlock constrained compute, improve utilization of existing infrastructure, and achieve higher returns.

At the same time, market demand is pushing beyond GPUs alone. IDC expects global AI infrastructure spending to reach $497 billion in 2026, up about 56% year over year, and demand is no longer limited to GPU systems; it also includes orchestration platforms, data pipelines, and CPU-based inference clusters.

Signals that favor Arm are growing stronger in the hyperscale cloud market

Microsoft and Amazon Web Services are the most closely watched cases right now, but they are only part of a broader architectural migration.

Google Cloud’s Axion is a family of CPU products customized based on the Arm architecture. Axion not only serves as the general-purpose CPU underpinning its cloud infrastructure, but also acts as the AI head node CPU in the latest generation of TPU systems, providing orchestration, networking, and infrastructure services to support stable operation of hyperscale AI systems.

NVIDIA has made the same architectural choice in its AI factory solution. Vera is NVIDIA’s latest Arm-based CPU and will serve as the compute foundation for next-generation AI systems such as Rubin. In these rack-level deployments, the CPU is responsible for coordinating memory, network, and accelerator resources.

Taken together, Microsoft Cobalt, Amazon Graviton, Google Axion, and NVIDIA Vera all point to the same conclusion: as AI infrastructure moves from model inference to autonomous agent execution, Arm-based CPUs are gradually becoming the general-purpose computing foundation of modern AI data centers, taking on critical roles such as resource orchestration and system coordination.