NVIDIA MGX GB200 NVL72 Reference Server Platform

NVIDIA MGX GB200 NVL72 Reference Server Platform

Brand: NVIDIA | Category: Servers

SKU: NVID-MGXGB200NVL72 | Part #: MGX-GB200-NVL72 | MPN: MGX-GB200-NVL72

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the NVIDIA MGX GB200 NVL72 Reference Server Platform

The NVIDIA MGX GB200 NVL72 is a full-rack reference platform integrating 36 Grace Blackwell Superchips — each pairing one NVIDIA Grace CPU with two NVIDIA Blackwell B200 GPUs — into a coherent 72-GPU system interconnected by fifth-generation NVLink Switch fabric. The NVLink Switch provides 1.8 TB/s of all-to-all GPU bandwidth across the entire rack, effectively making all 72 GPUs appear as a single massively parallel processor. At full precision the system delivers 720 petaFLOPS of FP8 compute and 1.4 exaFLOPS of FP4 AI compute, enabling training and inference of frontier large language models with hundreds of billions to trillions of parameters without model parallelism bottlenecks at node boundaries.

The platform is built on the NVIDIA MGX modular server architecture, which standardizes mechanical, thermal, power, and management interfaces to allow system builders to produce validated derivatives while adhering to NVIDIA-defined reference baselines. The Grace CPU component of each Superchip features 72 Arm Neoverse V2 cores and is connected to its paired B200 GPUs via a 900 GB/s NVLink-C2C chip-to-chip interconnect, eliminating PCIe bottlenecks and enabling unified memory addressing across CPU and GPU memory. Each rack column incorporates direct liquid cooling (DLC) as the primary thermal solution, supporting facility inlet water temperatures up to 45 °C and enabling extremely dense deployment in modern hyperscale and enterprise data centers.

Designed explicitly for frontier AI training, large-scale inference serving, and high-performance computing, the GB200 NVL72 supports NVIDIA's full software ecosystem including CUDA, cuDNN, TensorRT, NEMO, and the NVIDIA AI Enterprise platform. The system connects to the external fabric via 400 Gb/s InfiniBand or Ethernet ports per NVLink Switch tray, enabling multi-rack scale-out for exascale AI clusters. Its unified memory architecture, extreme bisection bandwidth, and direct liquid cooling collectively position this platform as the reference infrastructure for organizations developing and deploying next-generation generative AI, scientific simulation, and sovereign AI workloads.

Ideal for

  • Pre-training and fine-tuning of trillion-parameter large language models (LLMs) and multimodal foundation models requiring exaFLOP-scale compute within a single rack domain
  • High-throughput AI inference serving for LLMs where NVLink-unified GPU memory eliminates model-sharding latency and maximizes tokens-per-second per rack
  • Sovereign AI infrastructure deployments where governments and national research institutions require on-premises frontier model training capabilities at hyperscale density
  • Molecular dynamics simulation, quantum chemistry, and climate modeling workloads that benefit from the platform's unified 13.5 TB of HBM3e memory addressable across all 72 GPUs
  • Recommender system training and ranking model development at internet scale within cloud service provider GPU clusters built from multiple NVL72 racks
  • Drug discovery and genomics research pipelines combining HPC simulation and deep learning inference on a single unified rack-scale accelerator platform

Technical specifications

ManufacturerNVIDIA
Platform ArchitectureNVIDIA MGX GB200 NVL72 — rack-scale reference design
GPU Count per Rack72× NVIDIA Blackwell B200 GPUs
CPU Count per Rack36× NVIDIA Grace CPUs (72 Arm Neoverse V2 cores each)
Superchip Configuration36× Grace Blackwell Superchips (1× Grace + 2× B200 per Superchip)
AI Compute (FP4)1.4 ExaFLOPS
AI Compute (FP8)720 PetaFLOPS
HBM3e Memory per GPU192 GB HBM3e
Total HBM3e Memory (Rack)13,824 GB (≈ 13.5 TB) aggregate GPU memory
GPU Memory Bandwidth per GPU8 TB/s
CPU-to-GPU InterconnectNVLink-C2C at 900 GB/s bidirectional per Superchip
NVLink Switch All-to-All Bandwidth1.8 TB/s bisection bandwidth across all 72 GPUs
NVLink Generation5th Generation NVLink with NVLink Switch System
External Network ConnectivityUp to 400 Gb/s InfiniBand NDR or 400 Gb/s Ethernet per NVLink Switch tray
System Memory (CPU LPDDR5X)480 GB LPDDR5X per Grace CPU node (total ~17.3 TB across 36 nodes)
Storage InterfaceNVMe SSDs via PCIe Gen 5 on each compute tray; network-attached storage integration via external fabric
Thermal SolutionDirect Liquid Cooling (DLC), primary cooling; max facility inlet water temperature 45 °C
Rack Power Draw (TDP)~120 kW total rack power (reference configuration)
Power DistributionRack-integrated high-voltage DC (HVDC) or AC power shelf architecture per MGX spec
Form FactorFull-rack (42U–56U depending on build); NVIDIA MGX standardized mechanical envelope
ManagementBMC-based out-of-band management per node; NVIDIA Base Command Manager; DCGM GPU telemetry
Software EcosystemCUDA, cuDNN, TensorRT, NEMO Framework, NVIDIA AI Enterprise, Triton Inference Server
LaunchQ1 2025

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandNVIDIA
CategoryServers
SKUNVID-MGXGB200NVL72
Part NumberMGX-GB200-NVL72
ConditionNew
Platform ArchitectureNVIDIA MGX GB200 NVL72 — rack-scale reference design
GPU Count per Rack72× NVIDIA Blackwell B200 GPUs
CPU Count per Rack36× NVIDIA Grace CPUs (72 Arm Neoverse V2 cores each)
Superchip Configuration36× Grace Blackwell Superchips (1× Grace + 2× B200 per Superchip)
AI Compute (FP4)1.4 ExaFLOPS
AI Compute (FP8)720 PetaFLOPS
HBM3e Memory per GPU192 GB HBM3e
Total HBM3e Memory (Rack)13,824 GB (≈ 13.5 TB) aggregate GPU memory
GPU Memory Bandwidth per GPU8 TB/s
CPU-to-GPU InterconnectNVLink-C2C at 900 GB/s bidirectional per Superchip
NVLink Switch All-to-All Bandwidth1.8 TB/s bisection bandwidth across all 72 GPUs
NVLink Generation5th Generation NVLink with NVLink Switch System
External Network ConnectivityUp to 400 Gb/s InfiniBand NDR or 400 Gb/s Ethernet per NVLink Switch tray
System Memory (CPU LPDDR5X)480 GB LPDDR5X per Grace CPU node (total ~17.3 TB across 36 nodes)
Storage InterfaceNVMe SSDs via PCIe Gen 5 on each compute tray; network-attached storage integration via external fabric
Thermal SolutionDirect Liquid Cooling (DLC), primary cooling; max facility inlet water temperature 45 °C
Rack Power Draw (TDP)~120 kW total rack power (reference configuration)
Power DistributionRack-integrated high-voltage DC (HVDC) or AC power shelf architecture per MGX spec
Form FactorFull-rack (42U–56U depending on OEM build); NVIDIA MGX standardized mechanical envelope
ManagementBMC-based out-of-band management per node; NVIDIA Base Command Manager; DCGM GPU telemetry
Software EcosystemCUDA, cuDNN, TensorRT, NEMO Framework, NVIDIA AI Enterprise, Triton Inference Server
LaunchQ1 2025

Frequently Asked Questions about NVIDIA MGX GB200 NVL72 Reference Server Platform

What does the NVIDIA MGX GB200 NVL72 Reference Server Platform do?

The NVIDIA MGX GB200 NVL72 Reference Server Platform is built for enterprise data-center workloads — virtualization (VMware, Proxmox, Nutanix), private cloud, database hosting, and AI/ML training. It fits standard EIA-310 server racks and supports redundant PSUs and hot-swap drives common in production environments.

What are the headline specs of the NVIDIA MGX GB200 NVL72 Reference Server Platform?

Key specifications for the NVIDIA MGX GB200 NVL72 Reference Server Platform: new condition; manufacturer NVIDIA; platform architecture NVIDIA MGX GB200 NVL72 — rack-scale reference design; gpu count per rack 72× NVIDIA Blackwell B200 GPUs; cpu count per rack 36× NVIDIA Grace CPUs (72 Arm Neoverse V2 cores each); superchip configuration 36× Grace Blackwell Superchips (1× Grace + 2× B200 per Superchip); ai compute (fp4) 1.4 ExaFLOPS. Manufacturer part number MGX-GB200-NVL72. For the full datasheet with electrical, environmental, and compliance details, contact our pre-sales engineering team.