AMD Helios Is Its First Real Rack-Scale Answer to NVIDIA

AMD's MI455X and 72-GPU Helios rack finally challenge NVIDIA at the system level, but peak specifications are only the start of the comparison.

AMD Helios and NVIDIA Vera Rubin rack-scale systems compared by GPU count and aggregate HBM4 capacity

AMD used its Advancing AI event on July 23, 2026 to make a more serious claim than simply launching another fast accelerator. The new Instinct MI455X is designed for Helios, a 72-GPU rack-scale system that competes with NVIDIA's Vera Rubin NVL72 at the level where the largest AI systems are now built.

That distinction matters. NVIDIA's advantage is no longer just a good GPU or a mature software stack. It sells a complete computing system in which accelerators, CPUs, memory, switches, networking and software are designed together. Helios is AMD's clearest attempt yet to answer the whole system.

What AMD actually announced

The official MI455X specifications, published July 23, describe a 320-billion-transistor accelerator using AMD's fifth-generation CDNA architecture. Its compute dies use TSMC's 2 nm process, while other parts use 3 nm.

Each MI455X has 432 GB of HBM4 and 23.3 TB/s of peak memory bandwidth. Helios connects 72 accelerators into one scale-up domain, giving a full rack approximately 31.1 TB of GPU memory.

Specification AMD MI455X NVIDIA Rubin GPU
GPU memory 432 GB HBM4 288 GB HBM4
Peak memory bandwidth 23.3 TB/s 22 TB/s
Advertised low-precision compute 40.3 PFLOPS OCP MXFP4 50 PFLOPS NVFP4
Transistors 320 billion 336 billion
GPUs per rack-scale system 72 72

NVIDIA's figures come from its current Vera Rubin NVL72 specifications. Both companies label these as peak or preliminary values, and their low-precision formats are not identical enough to treat the table as a benchmark.

On paper, AMD's most obvious lead is memory capacity. Helios carries roughly 50 percent more aggregate HBM4 than the 20.7 TB listed for Vera Rubin NVL72.

That can matter for very large models, long contexts and inference systems that need to keep more parameters or key-value cache close to the accelerators. More local memory can reduce model partitioning and avoid some expensive movement between racks.

Capacity alone does not guarantee faster training or cheaper inference. The software must place data effectively, the interconnect must keep the GPUs fed, and applications must reach a useful fraction of the advertised bandwidth.

NVIDIA is already on its third generation of the NVL72 rack design. Rubin combines 72 GPUs with 36 Vera CPUs, NVLink 6 switching, ConnectX networking and BlueField DPUs. That installed design experience is difficult to express in a PFLOPS table.

AMD must prove that Helios can be manufactured, deployed and serviced reliably at scale. It must also demonstrate that ROCm and its communications libraries can turn the hardware into consistent application performance across real models.

This is the central test. A rack that looks competitive in a presentation can still lose on cluster setup time, software compatibility, failure recovery or achieved tokens per watt.

What AMD still has to prove

The important news is not that AMD has already beaten NVIDIA. No independent production testing supports that conclusion.

The important news is that AMD now has a credible architectural answer to NVIDIA's rack-scale strategy. The MI455X offers competitive per-GPU bandwidth, substantially more HBM capacity and a 72-accelerator domain rather than a collection of loosely connected cards. Tom's Hardware's July 23 analysis reaches a similar conclusion: this is AMD's strongest system-level challenge so far.

Four things will decide whether Helios becomes real competition:

  1. Independent training and inference results using the same models and precision.
  2. Power consumption measured for the complete rack, not just individual accelerators.
  3. ROCm stability and deployment effort at 72-GPU scale.
  4. Actual delivery volume and customer availability in the second half of 2026.

Until those results appear, Helios should be treated as credible hardware with unproven production economics. That is still a meaningful change. For the first time, AMD is challenging the shape of NVIDIA's AI system, not merely one chip inside it.