AMD Helios rack-scale accelerator: 31TB of HBM4 stacked in a single data center cabinet

AMD Helios Ships: The $1.2B Vultr Order That Puts 31TB of HBM4 in One Rack

AMD Helios Ships: Vultr’s $1.2B First Rack Order Puts 31TB of HBM4 on the Grid

On September 30, HPE announced a $1.2 billion order from Vultr to deploy AMD Helios AI Rack systems across Vultr’s U.S. data centers — the first customer order for the rack-scale design AMD unveiled in production form at its Advancing AI 2026 event in San Francisco on July 23. HPE stock jumped 5% on the news, and AMD closed at an all-time high of $633.91 on October 2, passing $1 trillion in market value. TokenPost reported on October 7 that Helios has moved into production, with customer shipments beginning in the second half of 2026.

That is the story most coverage framed as a chip race. It is not a chip story. Helios is the first time a GPU vendor sells the rack as the unit of compute, and the first time an open, Ethernet-based scale-up fabric has been ordered in volume against NVIDIA’s proprietary NVLink. For anyone who runs AI workloads — renting GPUs, deploying agents, fitting a large model into memory — the rack-as-product changes what you can rent, what it costs, and how much memory a single machine can hold.

What Helios Actually Is

Helios is a liquid-cooled, double-wide Open Compute Project Open Rack Wide chassis containing:

  • 72 AMD Instinct MI455X GPUs — CDNA 5 architecture, TSMC 2nm compute dies plus 3nm cache and I/O dies, roughly 320 billion transistors across 24 chiplets, 432GB of HBM4 per GPU, up to 23.3 TB/s of peak memory bandwidth per GPU, 40 PFLOPS of OCP MXFP4 and 20 PFLOPS of FP8 compute.
  • 18 sixth-generation EPYC “Venice” CPUs — the industry’s first x86 server processor in volume production on TSMC’s 2nm node, scaling to 256 Zen 6 cores, 512 threads, 16 memory channels, 1GB of L3 cache per socket, PCIe 6.0 and CXL 3.1.
  • AMD Pensando “Vulcano” AI NICs at 800 Gbps, with UALink over Ethernet (UALoE) for direct CPU-to-GPU connectivity, plus Pensando Salina DPUs for front-end offload.
  • AMD ROCm as the software layer.

Physically, 18 compute trays each carry four MI455X modules and one Venice host CPU, with two Vulcano NIC boards per tray — up to 2.4 Tbps of scale-up bandwidth per GPU. Four scale-up cartridges stitch all 72 GPUs into a single UALoE domain.

The rack-level numbers AMD publishes:

Metric Per rack
FP4 compute (OCP MXFP4) 2.9 exaflops
FP8 compute 1.4 exaflops
Aggregate HBM4 capacity 31 TB
Aggregate HBM bandwidth 1.7 PB/s
Scale-up bandwidth (inside the rack) 260 TB/s
Scale-out bandwidth (leaving the rack) 43 TB/s

The 31TB figure is the one that matters for practitioners. A single rack holds the weights of a multi-hundred-billion-parameter model plus a substantial inference cache, in memory, on one system — no tensor-parallel sharding across multiple machines, no cross-node traffic for the model itself. AMD rates the MI455X at up to 40 PFLOPS FP4 per chip, and a full rack at 2.9 exaflops. Independent benchmark history supports the trajectory: AMD’s Instinct MI350 series set a record 5.75M tokens/sec at 512-GPU scale in MLPerf Inference v6.1, and the MI355X posted results within single-digit percentage points of NVIDIA’s B200 in MLPerf Inference 6.0 (published April 1, 2026).

The Vultr Order Is the First Proof Point

Announcements of gigawatt commitments are easy. Orders with money attached are not. The Vultr deal, announced at HPE’s networking investor day, is HPE’s first Helios order, and it bundles HPE’s purpose-built networking — six HPE Juniper Networking QFX5252 scale-up Ethernet switch trays per rack — plus liquid cooling and services. HPE and AMD lined up this arrangement in December 2025, with HPE supplying the scale-up switching and software inside each rack.

Vultr, which HPE describes as the world’s largest privately held cloud infrastructure company, announced in July that it was taking pre-orders for reserved Helios capacity spanning 2027 through 2028. The HPE order puts HPE-built racks behind that program. Vultr CEO J.J. Kardwell framed it as speed to deployment: “Demand for high-performance AI infrastructure continues to outpace available capacity, and our work with AMD and HPE enables us to bring Helios rack-scale compute and scale-up networking online for customers faster.”

Estimated pricing for a full Helios rack runs around $5.25 million. That is not a purchase you make; it is capacity you rent, and the first parties with the balance sheet to rent it are already lined up.

The Demand Book Behind the Rack

Helios ships into a queue, not a market:

  • OpenAI — committed to a 6-gigawatt, multi-generation buildout centered on the MI450 platform (announced October 6, 2025), first gigawatt beginning in the second half of 2026, Helios online in Q4 2026. AMD issued OpenAI a warrant for up to 160 million AMD shares, vesting in tranches tied to deployment milestones.
  • Meta — 6 gigawatts of AMD Instinct across multiple chip generations (announced February 24, 2026), first 1GW of shipments in the second half of 2026.
  • Oracle — the most specific volume commitment: 50,000 AMD Instinct MI450-series GPUs in Helios-architecture superclusters, starting calendar Q3 2026, expanding in 2027. Oracle Cloud Infrastructure is named as the first hyperscaler to offer a publicly available AI supercluster on these chips.
  • Microsoft — announced July 20, 2026 that Azure deploys Helios at scale, with three new VM families: ND MI455X v7 for production-scale reasoning, search, and agentic inference; HDv2 (EPYC Venice, roughly 500 physical cores and 4TB of RAM) targeting agentic AI orchestration; and HXv2 for EDA and scientific computing. Pricing and launch regions were not announced.
  • Anthropic — up to 2 gigawatts of MI455X in Helios systems beginning early 2027, with AMD investing up to $5 billion in the Claude maker tied to milestones.

Twelve gigawatts of committed AMD capacity from OpenAI and Meta alone is the headline number. The structural point is narrower and more useful: for the first time, a hyperscaler AI buildout can be architected around someone other than NVIDIA, on a 2026 deployment timeline. NVIDIA still controls roughly 80–90% of the data-center GPU market. Default no longer means uncontested.

Why Open Standards Matter More Than Exaflops

NVIDIA’s rack-scale dominance is not built on GPU silicon. It is built on NVLink — a proprietary scale-up fabric you cannot buy from a second vendor. Helios runs scale-up on UALink over Ethernet inside an OCP Open Rack Wide chassis, with Ethernet switching from HPE/Juniper, and AMD co-authored the open standard Microsoft deployed Helios on. That means the switching, cabling, cooling, and service layers become commodity, sourced from a supply chain that is not owned by the GPU vendor.

For buyers, that is the price lever. Buyers with GPU-shortage scars from the past two years want two suppliers, not a hedge — the ability to play them off on price and availability. AMD’s own analyst-day projection is a $1 trillion market for its data-center chips by 2030, with the data-center segment growing at roughly 60% annually. The open fabric is what makes that competition real rather than nominal.

ROCm: The Software Wall Came Down First

Rack hardware is worthless if the software stack does not run your model. The material change over the past 18 months:

  • ROCm 7.2 (March 2026) is the first AMD release that reaches Ollama, LM Studio, llama.cpp, and vLLM parity with CUDA out of the box. FlashAttention-2 kernels for ROCm have been merged for both RDNA and CDNA targets, removing the long-standing long-context penalty.
  • PyTorch, vLLM, and SGLang all have official ROCm support in 2026. On MI300X and MI355X, ROCm reaches 90–95% of H100 throughput for standard LLM inference, and within 5–10% on standard PyTorch/vLLM serving.
  • ROCm 10 added AMD Infinity Storage with hipFILE features that move data straight from storage to GPU instead of through CPU RAM, NUMA-aware allocation, WSL2 support, and an agentic tooling layer (ROCm CLI, AMD Skills), riding a six-week release cadence.
  • A preview of ROCm.AI on ROCm 7.2.2 delivered a median 3.3x higher inference throughput across GLM-5, Kimi-K2.5, and DeepSeek-R1-0528 on 8x MI355X versus ROCm 7.0. ROCm Hyperloom reported a median 1.73x inference speedup in unattended evaluation, with gains from 1.35x to 7.31x.

The honest gap: workloads that depend on CUDA-only libraries — TensorRT-LLM, FlashAttention 3, NVIDIA NIM containers, custom PTX kernels — have no full ROCm equivalents, and Phoronix’s benchmark sweep puts ROCm 10–25% behind CUDA on equivalent silicon at the same precision. If your pipeline is PyTorch plus vLLM or SGLang, parity is close. If it is TensorRT-LLM, stay on CUDA.

What This Means for People Who Run AI Themselves

You will not rent a Helios rack. The consequences land on you anyway, through three channels.

Rented capacity gets cheaper and more plentiful

Helios capacity arrives on Azure (ND MI455X v7), OCI (50,000-GPU supercluster from Q3 2026), and Vultr. Second-supplier competition on price is the stated reason hyperscalers signed. Long-context agentic workloads — the ones where prefill dominates and KV cache pressure kills throughput — are exactly the workloads 432GB of HBM4 per GPU and 23.3 TB/s of bandwidth address. Memory-bound inference is where AMD’s capacity-per-dollar advantage wins on real workloads, not on paper benchmarks.

Memory bandwidth is the constraint, not parameter count

The Helios math scales down to your desk with the same physics. AMD’s consumer-scale version of the unified-memory idea is the Ryzen AI Max+ 395 (“Strix Halo”): 16 Zen 5 cores, a Radeon 8060S iGPU with 40 RDNA 3.5 compute units, an XDNA 2 NPU rated above 50 TOPS, and up to 128GB of shared LPDDR5X, with up to 96GB allocatable to the GPU. It runs Llama 3.3 70B in BF16 without sharding at around 14 tokens/sec, and Qwen3-235B at roughly 11 tokens/sec on a GMKtec EVO-X2. The caveat that matters: unified memory runs at about 256 GB/s, against 1,000+ GB/s on a discrete GPU. You buy capacity; you pay for bandwidth.

The procurement unit moved up a level

When the product is a rack, capacity is sold as reserved rack-scale blocks, and availability follows deployment schedules, not spot markets. Vultr’s program spans 2027–2028. Microsoft has not published a customer-access timetable. If you are planning a long-horizon inference budget, the second supplier existing is the useful fact; the capacity you can rent today is still the MI350-series generation.

Gear That Brings the Rack-Scale Idea to Your Desk

If the point of Helios is a model resident in memory on one machine, the desk-scale version of that argument is unified memory. The GMKtec EVO-X2 with 128GB LPDDR5X-8000 and a 2TB SSD is the cheapest in-stock way into a 128GB Strix Halo machine — $3,649.99 for the 128GB/2TB configuration, with a 64GB/1TB tier at $2,199.99 if you only need to fit models up to roughly 70B. The same listing carries the 64GB, 96GB, and 128GB tiers as variants of one product, so pick the memory size, not the ASIN.

For a rackable, managed version of the same architecture, the HP Z2 Mini G1a workstation ships an AMD Ryzen AI Max+ PRO 395 with 128GB of LPDDR5X-8533 (ECC), up to 96GB assignable to the Radeon 8060S, and an internal power supply that lets five units fit across a 4U rack — the small-form-factor echo of the rack-as-unit-of-compute idea, with Windows 11 Pro, WSL2, or Linux support.

If you are building a CPU-first agent box instead of a GPU box, the AMD Ryzen 9 9950X (16 cores / 32 threads, Zen 5, 5.7 GHz max boost, AM5, DDR5-5600) is the consumer-scale relative of the Venice server line, listed at $623.28. Pair it with a G.SKILL Trident Z5 Neo 64GB (2x32GB) DDR5-6000 AMD EXPO kit — EXPO is the AMD platform’s answer to XMP, and 64GB is enough to run a 2B decision model plus a mid-size model plus tooling on one box. Note that AM5 desktop memory tops out at 192GB across four DIMMs; you cannot reach Helios-style capacity on a desktop part, which is the honest limit of the desk-scale analogy.

What This Means

Helios entering production with a funded first order is the moment AMD’s rack-scale bet stops being a roadmap. Three things follow: procurement moved from GPUs to racks, the scale-up fabric is now an open standard with a real vendor supply chain, and ROCm parity means the software excuse for a single-supplier stack is gone for mainstream serving. For readers who run AI locally, the practical read is bandwidth, not exaflops — the same constraint that makes 432GB of HBM4 valuable in a rack is the constraint that makes 128GB of unified memory valuable on a desk, and the same competition that should push rented inference prices down over the next 12 months.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *