Supermicro GB200 NVL72 Rack-Scale System (SRS-GB200NVL72)

Supermicro GB200 NVL72 Rack-Scale System (SRS-GB200NVL72)

Brand: Supermicro | Category: Servers

SKU: SUPE-SRSGB200NVL7201 | Part #: SRS-GB200NVL72-01 | MPN: SRS-GB200NVL72-01

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the Supermicro GB200 NVL72 Rack-Scale System (SRS-GB200NVL72)

The Supermicro SRS-GB200NVL72 is a purpose-built rack-scale AI infrastructure system based on NVIDIA's Grace Blackwell Superchip architecture, housing 72 NVIDIA Blackwell GPUs and 36 NVIDIA Grace ARM-based CPUs within a single NVLink-connected fabric. The system is engineered as a unified compute domain rather than a collection of discrete servers, enabling all 72 GPUs to operate as one cohesive accelerated compute engine with fifth-generation NVLink providing 1.8 TB/s bidirectional bandwidth per GPU. This architecture eliminates traditional PCIe bottlenecks and allows the full rack to function as a single logical unit for the largest generative AI training and inference workloads.

The system leverages NVIDIA's NVLink Switch technology to create a fully non-blocking, all-to-all GPU interconnect across all 72 GPUs in the rack. Each GB200 NVL72 module pairs two Blackwell B200 GPUs with one Grace CPU via chip-to-chip NVLink, delivering 900 GB/s CPU-to-GPU bandwidth. The combined rack delivers a peak FP8 compute density exceeding 11.5 exaFLOPS, with 13.5 TB of aggregate HBM3e memory providing the capacity required for training and serving frontier-scale large language models with hundreds of billions to trillions of parameters.

Supermicro's implementation incorporates a direct liquid cooling architecture that handles the 120 kW+ rack-level thermal load with precision manifold distribution, ensuring sustained GPU boost performance across all 72 accelerators simultaneously. The system ships as a pre-integrated, factory-validated rack unit designed to minimize on-site integration complexity for hyperscale and enterprise AI data centers, supporting direct integration with high-performance storage fabrics, InfiniBand or Ethernet networking, and Supermicro's full lifecycle management tooling via IPMI and Redfish-compliant BMC interfaces.

Ideal for

  • Training frontier large language models with 100B+ parameters where GPU-to-GPU bandwidth and memory capacity are the primary scaling constraints
  • High-throughput generative AI inference serving for production LLM deployments requiring maximum tokens-per-second density per rack
  • Multimodal foundation model development combining vision, language, and speech modalities at scale within a single unified compute domain
  • AI supercomputing cluster construction where multiple GB200 NVL72 racks are interconnected via InfiniBand NDR400 for petascale to exascale training jobs
  • Scientific simulation and digital twin workloads in climate modeling, drug discovery, and physics simulation that require tightly coupled GPU collaboration
  • Enterprise AI factory deployments consolidating multiple model training pipelines on a shared, high-utilization rack-scale infrastructure

Technical specifications

ManufacturerSupermicro
Product LineGrace Blackwell Rack-Scale System
ModelSRS-GB200NVL72
GPU ArchitectureNVIDIA Blackwell (B200)
Total GPUs per Rack72x NVIDIA B200 GPUs
CPU ArchitectureNVIDIA Grace (ARM Neoverse V2)
Total CPUs per Rack36x NVIDIA Grace CPUs (paired 2:1 GPU-to-CPU via NVLink C2C)
GPU InterconnectNVLink 5th Generation, all-to-all non-blocking fabric across all 72 GPUs
NVLink Bandwidth per GPU1.8 TB/s bidirectional
Grace-to-Blackwell C2C Bandwidth900 GB/s bidirectional per GB200 Superchip pair
Peak AI Compute (FP8)>11.5 ExaFLOPS aggregate across 72 B200 GPUs
Total HBM3e Capacity~13.5 TB aggregate (192 GB HBM3e per B200 GPU)
HBM3e Memory Bandwidth per GPU8 TB/s
Grace CPU MemoryLPDDR5X with 480 GB/s bandwidth per Grace CPU; 128 GB per CPU module
Thermal DesignDirect liquid cooling (DLC) with manifold distribution; supports >120 kW rack-level thermal load
Rack Form FactorFull rack (42U-class rack-scale unit), factory-integrated
Network Fabric SupportInfiniBand NDR400 and/or 400GbE per compute tray for scale-out interconnect
Storage InterfaceNVMe-oF / GPUDirect Storage compatible; external high-performance storage fabric integration
System ManagementIPMI 2.0, Redfish API, Supermicro BMC, NVIDIA DOCA integration
Software EcosystemNVIDIA AI Enterprise, CUDA 12.x, NVLink Switch System software, NGC container registry compatible
Power SupplyRedundant high-efficiency 277V AC or 380V DC power shelf; N+N redundancy configuration
Target DeploymentHyperscale AI data centers, enterprise AI factories, national AI supercomputing facilities
Launch2025 Q3

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandSupermicro
CategoryServers
SKUSUPE-SRSGB200NVL7201
Part NumberSRS-GB200NVL72-01
ConditionNew
Product LineGrace Blackwell Rack-Scale System
ModelSRS-GB200NVL72
GPU ArchitectureNVIDIA Blackwell (B200)
Total GPUs per Rack72x NVIDIA B200 GPUs
CPU ArchitectureNVIDIA Grace (ARM Neoverse V2)
Total CPUs per Rack36x NVIDIA Grace CPUs (paired 2:1 GPU-to-CPU via NVLink C2C)
GPU InterconnectNVLink 5th Generation, all-to-all non-blocking fabric across all 72 GPUs
NVLink Bandwidth per GPU1.8 TB/s bidirectional
Grace-to-Blackwell C2C Bandwidth900 GB/s bidirectional per GB200 Superchip pair
Peak AI Compute (FP8)>11.5 ExaFLOPS aggregate across 72 B200 GPUs
Total HBM3e Capacity~13.5 TB aggregate (192 GB HBM3e per B200 GPU)
HBM3e Memory Bandwidth per GPU8 TB/s
Grace CPU MemoryLPDDR5X with 480 GB/s bandwidth per Grace CPU; 128 GB per CPU module
Thermal DesignDirect liquid cooling (DLC) with manifold distribution; supports >120 kW rack-level thermal load
Rack Form FactorFull rack (42U-class rack-scale unit), factory-integrated
Network Fabric SupportInfiniBand NDR400 and/or 400GbE per compute tray for scale-out interconnect
Storage InterfaceNVMe-oF / GPUDirect Storage compatible; external high-performance storage fabric integration
System ManagementIPMI 2.0, Redfish API, Supermicro BMC, NVIDIA DOCA integration
Software EcosystemNVIDIA AI Enterprise, CUDA 12.x, NVLink Switch System software, NGC container registry compatible
Power SupplyRedundant high-efficiency 277V AC or 380V DC power shelf; N+N redundancy configuration
Target DeploymentHyperscale AI data centers, enterprise AI factories, national AI supercomputing facilities
Launch2025 Q3

Frequently Asked Questions about Supermicro GB200 NVL72 Rack-Scale System (SRS-GB200NVL72)

What does the Supermicro GB200 NVL72 Rack-Scale System (SRS-GB200NVL72) do?

The Supermicro GB200 NVL72 Rack-Scale System (SRS-GB200NVL72) is built for enterprise data-center workloads — virtualization (VMware, Proxmox, Nutanix), private cloud, database hosting, and AI/ML training. It fits standard EIA-310 server racks and supports redundant PSUs and hot-swap drives common in production environments.

What are the headline specs of the Supermicro GB200 NVL72 Rack-Scale System (SRS-GB200NVL72)?

Key specifications for the Supermicro GB200 NVL72 Rack-Scale System (SRS-GB200NVL72): new condition; manufacturer Supermicro; product line Grace Blackwell Rack-Scale System; model SRS-GB200NVL72; gpu architecture NVIDIA Blackwell (B200); total gpus per rack 72x NVIDIA B200 GPUs; cpu architecture NVIDIA Grace (ARM Neoverse V2). Manufacturer part number SRS-GB200NVL72-01. For the full datasheet with electrical, environmental, and compliance details, contact our pre-sales engineering team.