Join the 155,000+ IMP followers

electronics-journal.com

Rackscale Compute Platform for High-Density Inference Operations

Microsoft incorporates AMD Helios rackscale systems, EPYC Venice processors, and Pensando DPUs to scale artificial intelligence infrastructure on Azure.

  www.amd.com
Rackscale Compute Platform for High-Density Inference Operations

AMD and Microsoft Corporation have expanded their technical collaboration to integrate an open, rackscale computing platform into Microsoft Azure data centers. The deployment focuses on accelerating high-density artificial intelligence inference, data processing pipelines, and specialized engineering workloads across cloud infrastructure.

System Architecture and Integration Logic
The hardware and software integration addresses the thermal, interconnect, and memory bandwidth bottlenecks associated with large language model deployment. AMD supplies the full-stack architecture, while Microsoft manages cluster integration, virtual machine provisioning, and service orchestration within the Azure ecosystem.

The system relies on the AMD Helios Rackscale Solution, an integrated hardware architecture combining AMD Instinct MI455X graphics processing units, sixth-generation AMD EPYC Venice central processing units, and Pensando data processing units. System components communicate via an open software framework using the AMD ROCm open ecosystem, allowing direct memory access and low-latency interconnects across distributed node architectures.

Hardware Allocation and Network Acceleration
The deployment divides responsibilities between compute processing, virtualized infrastructure, and dedicated network offloading:
  • Inference Compute Infrastructure: AMD Helios units handle frontier model artificial intelligence inference, processing high-throughput matrix operations for multi-modal and generative workloads.
  • Virtual Machine Provisioning: Azure HDv2 virtual machines leverage sixth-generation EPYC processors to execute data preparation, search, and agentic workflows. Azure HXv2 virtual machines utilize the same processor architecture to process high-memory electronic design automation tasks for semiconductor engineering.
  • Packet Processing and Security: Network tasks are offloaded to AMD Pensando DPUs integrated with the Azure Boost system. This architecture offloads storage and networking stacks from host CPUs to specialized hardware, increasing connection throughput and reducing system overhead.
Field Implementation and Operational Characteristics
Implementation follows a phased hardware rollout across global Azure data centers, with initial shipments of the rackscale architecture scheduled for the second half of 2026. The integrated compute stacks interface directly with Azure Foundry Managed Compute, allowing enterprise users to deploy production models without manual hardware provisioning.

By offloading system telemetry, network routing, and packet filtering to specialized data processing units, the architecture maintains deterministic network latency under peak cluster load. The open architecture framework allows the system to scale across multi-rack topologies, providing predictable execution times for enterprise data pipelines and high-performance computing tasks.

Edited by Evgeny Churilov, Induportals Media - Adapted by AI.

www.amd.com

  Ask For More Information…

LinkedIn
Pinterest

Join the 155,000+ IMP followers