CNBC

AMD's Helios Bets the AI Market Is Too Big to Be Nvidia's Alone

Forrest Norrod & Vamsi Boppana· Data center and AI leaders at AMD at AMD
·~18 min·English·CNBC
GPUAI InfrastructureInferenceBusiness Strategy
TL;DR

AMD unveils Helios, its first rack-scale AI system, as an open, full-stack challenge to Nvidia, betting that a market Nvidia controls more than 95% of is expanding fast enough for a real second source to win.

01The Stakes

Helios Is AMD's Make-or-Break Bet

Nvidia controls more than 95% of the data-center GPU market, and Helios is AMD's make-or-break bid to gain real share, with outside analysts sketching a path to 20 to 25%.

So it's absolutely our aspiration to be able to gain market share.

AMD, CNBC
Key Insight
The 95%-to-single-digit gap is why the sharper question isn't whether Helios is faster, but whether demand so outstrips supply that AMD gets bought simply for existing. AMD's own goal is stated only as gaining share, while an outside analyst sketches a path to 20 to 25%, either way a bid for a structural second-source role, not a single-digit niche.

02The Machine

Helios Fuses Four AMD Businesses Into One Rack

Helios packs 72 GPUs and 18 CPUs into one liquid-cooled rack and pulls together the four things AMD builds in-house: GPUs, CPUs, networking, and software.

It's our baby. It's 72 GPUs and 18 CPUs per rack.

Forrest Norrod, CNBC
Key Insight
Building all four layers in-house is the strategic bet: it lets AMD co-optimize the rack the way Nvidia does, tuning GPU, EPYC CPU, networking, and ROCm software as one system, so integration, not any single chip, becomes the product.

03The CPU Argument

Agentic AI Leans on the CPU

As call-and-answer chatbots give way to heavily orchestrated agentic AI, the general-compute CPU that feeds and schedules the GPUs turns AMD's oldest strength into a differentiator.

EPYC is clearly a leadership CPU. It's been leading in performance and power efficiency for most of the last decade, and I think it's critical in terms of feeding the GPU and keeping the whole system orchestrated and running efficiently.

Forrest Norrod, CNBC
Key Insight
This is AMD's most under-appreciated argument: agentic workloads orchestrate many tool calls and models, leaning on the CPU that schedules and feeds the GPUs, which is the exact part, EPYC, where AMD has led on performance and power efficiency for most of the last decade.

04Open vs Walled Garden

AMD's Answer to CUDA Is to Go Open

AMD's wedge against Nvidia's proprietary CUDA is openness, from open standards and the ROCm stack to first-class support for PyTorch, vLLM, and SGLang, even as CUDA's ecosystem stays far ahead.

We support the open frameworks. So one of these large customers is a big PyTorch house. Another one is a big proponent of vLLM. Another one is SGLang. So the fact that these open-source communities are well supported, I think really is valuable to them.

Vamsi Boppana, CNBC
Key Insight
Open is both a genuine architectural choice and a competitive necessity: AMD can't out-CUDA CUDA, so it reframes Nvidia's proprietary full-stack lead as a walled garden and courts the open-source serving communities, vLLM and SGLang, that already run frontier inference.

05Execution

You Can't Suddenly Build Helios

A first-generation system is a leap of faith, so AMD's answer is the three-generation EPYC roadmap it says it delivered exactly, plus the acquisitions, Pensando, Xilinx, and ZT Systems, that bought the missing pieces of the stack.

We couldn't have suddenly said that, hey, let's build Helios without having all of the investments there.

Vamsi Boppana, CNBC
Key Insight
The subtext is that a rack-scale system is an integration problem, not a chip problem: AMD spent years and tens of billions acquiring the networking, AI, and server-build capabilities before the GPU could headline, and it points to the three-generation EPYC roadmap it says it delivered exactly to argue the first generation is a safe bet.

06The Economics

The End of Token-Maxing

AMD is selling the lowest total cost of ownership and cost per token as the industry shifts from token-maxing toward practical token utilization that companies can actually justify.

We're seeing a little bit of a fall off of the token maxing era, and we're starting to see things like, hey, how do we get to practical token utilization to bring costs for intelligence to a level where companies can justify the spend?

Forrest Norrod, CNBC
Key Insight
Token-maxing fading matters because it moves the goalposts from peak throughput to cost per useful token, a memory-bandwidth and total-cost game where AMD's higher-memory, full-stack rack is designed to compete on the roofline, not just raw FLOPs.

07The Market

Not a Zero-Sum Game

AMD frames the fight as non-zero-sum: the market is expanding so fast, and capacity so constrained, that Microsoft, Meta, OpenAI, and Oracle can commit to Helios without Nvidia having to lose.

First of all, we don't believe that, you know, this is like a zero sum game at all. Right? You know the market is expanding. And this is an incredible, incredible growth cycle.

AMD, CNBC
Key Insight
The non-zero-sum framing is convenient but not baseless: with capacity the binding constraint, a credible second supplier de-risks every hyperscaler's buildout, which is why Microsoft, Meta, OpenAI, and Oracle would rather AMD succeed than depend on a single vendor.

08The Bottleneck

The Race Is for Everything but the Chip

The real bottleneck is everything but the chip, from power and water to wafer capacity, advanced packaging, and HBM memory, and AMD is answering each with efficiency programs and supply commitments.

it's up to 432GB of HBM memory per GPU. We have close relationships with all three of the major memory suppliers, so we've been able to secure all the memory we need.

Forrest Norrod, CNBC
Key Insight
Notice the pattern in AMD's answers: efficiency programs for power, TSMC and Arizona for wafers, ten billion dollars to ASE for packaging, deals with all three memory makers for HBM. The competitive moat in this cycle is supply-chain lockup as much as silicon.