NVIDIA GB200 NVL72 Rack-Scale System

NVIDIA GB200 NVL72 Rack-Scale System

Brand: NVIDIA | Category: GPUs

SKU: 900-21761-0040-000 | Part #: 900-21761-0040-000 | MPN: 900-21761-0040-000

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the NVIDIA GB200 NVL72 Rack-Scale System

The NVIDIA GB200 NVL72 is a rack-scale AI supercomputing system built on the Blackwell Ultra architecture, integrating 36 GB200 Superchips into a single liquid-cooled rack unit. Each GB200 Superchip pairs two NVIDIA B200 Tensor Core GPUs with one NVIDIA Grace CPU via NVLink-C2C, delivering a tightly unified compute fabric across the entire rack. The system interconnects all 72 B200 GPUs using fifth-generation NVLink switching, creating a single logical GPU with 1.4 TB of aggregate HBM3e memory and 130 TB/s of all-to-all GPU bandwidth — enabling model parallelism at a scale previously requiring multi-rack configurations.

Designed specifically for the demands of large language model training and inference, the GB200 NVL72 delivers up to 1.4 ExaFLOPS of AI performance at FP4 precision and 720 PFLOPS at FP8, with full support for FP16, BF16, TF32, and FP64 compute modes. The Grace CPUs within the system provide 144 high-performance Arm Neoverse V2 cores and up to 17.6 TB/s of CPU-to-GPU bandwidth via NVLink-C2C, eliminating traditional PCIe bottlenecks and enabling direct, high-bandwidth access to GPU memory from the host processing layer. The system ships as a complete rack-scale reference design inclusive of NVLink switches, high-speed networking interfaces, power infrastructure, and direct liquid cooling (DLC) subsystems.

The GB200 NVL72 is purpose-built for hyperscale data centers and enterprise AI infrastructure teams deploying frontier-scale workloads including trillion-parameter model training, real-time inference serving at massive throughput, and advanced scientific simulation. It is the foundational compute platform for customers building sovereign AI infrastructure, next-generation cloud services, and high-performance computing (HPC) clusters demanding maximum GPU-to-GPU communication bandwidth in a rack-contained form factor.

Ideal for

  • Training and fine-tuning of trillion-parameter large language models and multimodal foundation models requiring high-bandwidth GPU collective communication
  • High-throughput generative AI inference serving, supporting simultaneous large-batch requests with low latency for enterprise and cloud deployments
  • Accelerated scientific simulation and HPC workloads including molecular dynamics, climate modeling, and computational fluid dynamics at ExaFLOPS-scale
  • Sovereign AI data center builds for government and national research organizations requiring self-contained, rack-scale AI compute infrastructure
  • Large-scale recommendation system training and ranking model development for e-commerce, media, and advertising platforms
  • Drug discovery and genomics workloads leveraging FP64 precision and massive unified GPU memory for complex biomolecular simulations

Technical specifications

ManufacturerNVIDIA
Manufacturer Part Number900-21761-0040-000
Product FamilyBlackwell
GPU ArchitectureNVIDIA Blackwell
System Configuration72x NVIDIA B200 Tensor Core GPUs, 36x NVIDIA Grace CPUs (36 GB200 Superchips)
Total GPU Memory1.4 TB HBM3e
GPU Memory Per B200192 GB HBM3e
GPU Memory Bandwidth Per B2008 TB/s
Total AI Performance (FP4)1.4 ExaFLOPS
Total AI Performance (FP8)720 PFLOPS
GPU InterconnectNVLink 5 (fifth-generation), 130 TB/s all-to-all bisection bandwidth
CPU-to-GPU InterconnectNVLink-C2C, up to 900 GB/s per direction per Superchip
CPUNVIDIA Grace (Arm Neoverse V2), 72 cores per CPU, 144 CPU cores total per Superchip pair
CPU MemoryLPDDR5X per Grace CPU
Supported PrecisionsFP4, FP8, FP16, BF16, TF32, FP32, FP64, INT8
CoolingDirect Liquid Cooling (DLC)
Form FactorRack-scale system (full rack)
NetworkingNVIDIA ConnectX-7 / BlueField-3 high-speed networking interfaces
Operating System SupportLinux (Ubuntu, RHEL, and compatible enterprise distributions)
CUDA CompatibilityCUDA 12.x and later

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandNVIDIA
CategoryGPUs
SKU900-21761-0040-000
Part Number900-21761-0040-000
ConditionNew
Manufacturer Part Number900-21761-0040-000
Product FamilyBlackwell
GPU ArchitectureNVIDIA Blackwell
System Configuration72x NVIDIA B200 Tensor Core GPUs, 36x NVIDIA Grace CPUs (36 GB200 Superchips)
Total GPU Memory1.4 TB HBM3e
GPU Memory Per B200192 GB HBM3e
GPU Memory Bandwidth Per B2008 TB/s
Total AI Performance (FP4)1.4 ExaFLOPS
Total AI Performance (FP8)720 PFLOPS
GPU InterconnectNVLink 5 (fifth-generation), 130 TB/s all-to-all bisection bandwidth
CPU-to-GPU InterconnectNVLink-C2C, up to 900 GB/s per direction per Superchip
CPUNVIDIA Grace (Arm Neoverse V2), 72 cores per CPU, 144 CPU cores total per Superchip pair
CPU MemoryLPDDR5X per Grace CPU
Supported PrecisionsFP4, FP8, FP16, BF16, TF32, FP32, FP64, INT8
CoolingDirect Liquid Cooling (DLC)
Form FactorRack-scale system (full rack)
NetworkingNVIDIA ConnectX-7 / BlueField-3 high-speed networking interfaces
Operating System SupportLinux (Ubuntu, RHEL, and compatible enterprise distributions)
CUDA CompatibilityCUDA 12.x and later

Frequently Asked Questions about NVIDIA GB200 NVL72 Rack-Scale System

What server platforms accept the NVIDIA GB200 NVL72 Rack-Scale System?

Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.

How long is the lead time on AI GPUs?

Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.

Do you supply matched networking (Quantum InfiniBand / Spectrum-X)?

Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.

Can you help with NVIDIA AI Enterprise licensing?

Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.