AMD has expanded its AI infrastructure portfolio with the launch of Helios, an open, rackscale AI infrastructure designed for frontier AI and sovereign computing. Helios is built around AMD's next-generation Instinct GPUs, EPYC Venice processors, Pensando networking and the ROCm software stack.
The Significance of Helios in AMD's AI Strategy
For years, AMD has been a strong contender in the CPU market with its EPYC processors, and its Instinct GPUs have gradually carved out a niche in high-performance computing. However, the AI boom, dominated by Nvidia, has left AMD scrambling for a comprehensive solution. Helios represents AMD's most ambitious attempt yet to challenge Nvidia's dominance by offering a complete, integrated rack-scale system rather than individual components. This move mirrors Nvidia's own strategy with its DGX and Vera Rubin platforms, but with a key differentiator: openness.
According to Pareekh Jain, CEO at EIIRTrend & Pareekh Consulting, "Helios is AMD's first complete AI rack system with GPUs, CPUs, and networking built together, instead of selling separate chips. It is well suited for training large AI models, memory heavy models, long context processing and high volume inference, and AMD's biggest shot yet at challenging Nvidia's dominance."
The launch is particularly timely as enterprises and cloud providers seek alternatives to Nvidia amid supply constraints and rising costs. AMD has also secured an early hyperscale deployment for Helios with Microsoft agreeing to deploy it to power its frontier model AI inference, its AI customers, and support Azure AI services. This endorsement from a major cloud player adds credibility and could accelerate enterprise adoption.
The Architecture Behind Helios
Unlike previous AMD AI offerings centred on individual accelerators, Helios is designed as a complete rack-scale system integrating compute, networking and software. The AMD Helios rackscale design includes 72 AMD Instinct MI455X GPUs with AMD EPYC Venice CPUs and AMD Pensando Vulcano networking using UALink, optimized for compute, data movement, and system efficiency.
The platform supports both OCP and MX data types, delivering up to 2.9 EFLOPS of FP4 and 1.4 EFLOPS of FP8 compute for AI training and inference. This performance positions Helios to handle the most demanding AI workloads, including large language model training, real-time inference, and complex scientific simulations.
Memory is a standout feature. Helios integrates 31TB of HBM4 memory with 19.6TB/s of memory bandwidth, approximately 50% more total memory than Nvidia's competing system. This allows it to run very large AI models that otherwise would require multiple nodes. The system also employs a liquid-cooling design with quick-disconnect connections to efficiently dissipate heat, enabling denser rack configurations and lower operational costs.
Helios is built on open standards including OCP Open Rack Wide (ORW), Ultra Accelerator Link (UALink), and Ultra Ethernet Consortium (UEC). This contrasts with Nvidia's proprietary NVLink and InfiniBand technologies. Jain explains, "It uses open, industry-standard connections instead of Nvidia's private technology, giving buyers more flexibility."
On the security front, Helios incorporates a hardware root of trust and continuous attestation at every layer. It supports hardware-enforced isolation, encrypted memory and interconnects to help protect AI models, data and workloads in multi-tenant environments. This is critical as enterprises move sensitive AI workloads to shared cloud infrastructure.
The Software Challenge
While the launch of Helios might help AMD close the hardware gap with Nvidia's rack-scale systems, it will be the software compatibility that will be the real driver of enterprise adoption. Nvidia's CUDA ecosystem has a 15-20 year head start, and almost every AI tool, tutorial, and codebase defaults to it. AMD has been working to close this gap with its ROCm software platform, but the progress has been uneven.
AMD is expanding its ROCm AI software platform to support frameworks including PyTorch, TensorFlow, and JAX, for enabling high-throughput inference and efficient distributed training while preserving familiar developer workflows. However, Jain cautions, "Software has been AMD's weak spot. AMD has improved ROCm a lot but it still lags behind on the newest, most specialized optimizations, and setup is more complicated."
For everyday AI work, ROCm is usable but for cutting-edge performance, CUDA still leads. This means enterprises must carefully evaluate whether their AI models and tools are compatible with AMD's stack before committing. Some advanced tools remain CUDA-only, though the community and AMD are working to expand coverage.
Evaluating the Trade-offs for CIOs
For CIOs evaluating AI infrastructure, Helios launch brings in another option to a market that has largely revolved around Nvidia's dominance. But when considering Helios, CIOs will have to evaluate factors such as performance, software readiness, deployment models, procurement timelines and total cost of ownership before committing to a platform.
While AMD has not publicly announced a specific price tag for the Helios, Jain believes it to be noticeably cheaper to buy and run with lower chip prices and lower power use per GPU. This could translate into significant savings over the lifetime of a large AI cluster, especially as power costs continue to rise.
Jain notes, "It gives companies a real second option besides Nvidia, easing supply shortages and giving leverage in negotiations. The catch is software, where teams need to check whether their AI tools run well on AMD's stack, since some advanced tools are still CUDA only."
For CIOs planning to deploy both, Jain warns the two systems can't be plugged together into one combined machine as they use different, incompatible connection technology. But companies can and do run both side by side in the same data center, just as separate systems handling different jobs. This dual-vendor strategy is already being adopted by some hyperscalers and large enterprises to avoid vendor lock-in and optimize costs.
The market reaction to Helios will likely depend on how quickly AMD can address software gaps and provide seamless integration into existing AI workflows. With Microsoft as an early adopter, AMD has a strong reference case that could persuade other cloud providers and enterprises to experiment with Helios. However, Nvidia's entrenched position means AMD will need sustained investment in both hardware and software to gain meaningful market share.
In the broader context, Helios represents a maturing of AMD's AI offerings from chips to systems. It also underscores the industry's shift toward integrated, rack-scale solutions that promise better performance, power efficiency, and manageability. For CIOs, the arrival of a credible alternative to Nvidia's platforms is welcome news, potentially driving down costs and increasing innovation in AI infrastructure.
Source: Network World News