NVIDIA GB200 NVL72 Rack-Scale Compute Unit

NVIDIA GB200 NVL72 Rack-Scale Compute Unit

Brand: NVIDIA | Category: GPUs

SKU: 920-23687-2570-000 | Part #: 920-23687-2570-000 | MPN: 920-23687-2570-000

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the NVIDIA GB200 NVL72 Rack-Scale Compute Unit

The NVIDIA GB200 NVL72 is a rack-scale AI supercomputing system integrating 36 Grace CPUs and 72 Blackwell GPUs into a single NVLink-connected compute unit. Built on NVIDIA's fifth-generation NVLink fabric, the system delivers 1.4 exaflops of AI compute (FP4) and 720 petaflops of FP8 performance, enabling training and inference workloads at a scale previously achievable only by multi-rack HPC clusters. The 72 Blackwell B200 GPUs within the rack share a unified 8.6 TB aggregate HBM3e memory pool, eliminating traditional PCIe bottlenecks and allowing model parallelism across the entire rack as a single logical GPU domain.

The GB200 NVL72 is engineered for the most demanding large language model (LLM) training and inference workloads, including models exceeding one trillion parameters. Each B200 GPU incorporates the second-generation Transformer Engine with FP4 and FP8 mixed-precision support, dual-chip Blackwell GPU architecture per die, and fifth-generation NVLink providing 1.8 TB/s of bidirectional bandwidth per GPU. The 36 integrated Grace CPUs (Arm Neoverse V2 architecture) are connected to their paired B200 GPUs via NVLink-C2C, delivering 900 GB/s of coherent CPU-GPU interconnect bandwidth — eliminating the PCIe memory copy overhead that constrains conventional accelerated servers.

Designed for liquid-cooled data center deployments, the GB200 NVL72 utilizes a direct liquid cooling (DLC) architecture to manage the thermal demands of rack-scale AI compute. The system is well-suited for deployment in modern AI factories and hyperscale data centers equipped with rear-door or direct liquid cooling infrastructure. NVIDIA's NVLink Switch System within the rack enables all-to-all GPU communication at full bandwidth, making the GB200 NVL72 a foundational building block for sovereign AI infrastructure, cloud service provider GPU clusters, and enterprise-scale AI platforms across the UAE, GCC, EMEA, and APAC regions. Omnixon Global offers worldwide supply of this system to qualified enterprise, government, and research organizations.

Ideal for

  • Training and fine-tuning large language models and multimodal foundation models with parameter counts exceeding 100 billion, leveraging the unified 8.6 TB HBM3e memory pool and 1.4 exaflops of FP4 AI compute
  • High-throughput LLM inference serving for enterprise AI platforms requiring maximum token-per-second throughput with reduced per-token energy consumption compared to prior-generation GPU systems
  • Sovereign AI and national AI infrastructure deployments requiring rack-scale, self-contained AI supercomputing capability with predictable, high-bandwidth all-to-all GPU communication
  • Scientific simulation and digital twin workloads in energy, climate modeling, and life sciences that demand extreme floating-point throughput and large unified memory capacity
  • Cloud service provider GPU-as-a-Service (GPUaaS) buildouts targeting AI model training and inference services, where rack-level NVLink connectivity maximizes utilization and reduces multi-tenant overhead
  • Genomics, drug discovery, and computational biology research requiring simultaneous large-scale data ingestion and deep learning inference pipelines within a single coherent compute domain

Technical specifications

ManufacturerNVIDIA
Manufacturer Part Number920-23687-2570-000
Product FamilyGB200 NVL72
GPU ArchitectureNVIDIA Blackwell
GPU Count per Rack72 × NVIDIA B200 GPUs
CPU Count per Rack36 × NVIDIA Grace CPUs (Arm Neoverse V2)
AI Performance (FP4)1.4 ExaFLOPS per rack
AI Performance (FP8)720 PFLOPS per rack
GPU Memory per GPU192 GB HBM3e
Total Aggregate GPU Memory8.6 TB HBM3e (across 72 GPUs)
GPU Memory Bandwidth per GPU8 TB/s
CPU-GPU InterconnectNVLink-C2C at 900 GB/s bidirectional per Grace-Blackwell Superchip
NVLink GPU-to-GPU Bandwidth1.8 TB/s bidirectional per GPU
NVLink Generation5th Generation NVLink with NVLink Switch System
Interconnect TopologyNVLink Switch System enabling full all-to-all 72-GPU connectivity within the rack
Transformer Engine2nd Generation (FP4 and FP8 mixed precision support)
CoolingDirect Liquid Cooling (DLC)
Form FactorRack-Scale System (full rack)
Target InfrastructureLiquid-cooled AI data center / AI factory deployments
Supported Precision FormatsFP4, FP8, FP16, BF16, TF32, FP64
Product SegmentData Center / AI Infrastructure

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandNVIDIA
CategoryGPUs
SKU920-23687-2570-000
Part Number920-23687-2570-000
ConditionNew
Manufacturer Part Number920-23687-2570-000
Product FamilyGB200 NVL72
GPU ArchitectureNVIDIA Blackwell
GPU Count per Rack72 × NVIDIA B200 GPUs
CPU Count per Rack36 × NVIDIA Grace CPUs (Arm Neoverse V2)
AI Performance (FP4)1.4 ExaFLOPS per rack
AI Performance (FP8)720 PFLOPS per rack
GPU Memory per GPU192 GB HBM3e
Total Aggregate GPU Memory8.6 TB HBM3e (across 72 GPUs)
GPU Memory Bandwidth per GPU8 TB/s
CPU-GPU InterconnectNVLink-C2C at 900 GB/s bidirectional per Grace-Blackwell Superchip
NVLink GPU-to-GPU Bandwidth1.8 TB/s bidirectional per GPU
NVLink Generation5th Generation NVLink with NVLink Switch System
Interconnect TopologyNVLink Switch System enabling full all-to-all 72-GPU connectivity within the rack
Transformer Engine2nd Generation (FP4 and FP8 mixed precision support)
CoolingDirect Liquid Cooling (DLC)
Form FactorRack-Scale System (full rack)
Target InfrastructureLiquid-cooled AI data center / AI factory deployments
Supported Precision FormatsFP4, FP8, FP16, BF16, TF32, FP64
Product SegmentData Center / AI Infrastructure

Frequently Asked Questions about NVIDIA GB200 NVL72 Rack-Scale Compute Unit

What server platforms accept the NVIDIA GB200 NVL72 Rack-Scale Compute Unit?

Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.

How long is the lead time on AI GPUs?

Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.

Do you supply matched networking (Quantum InfiniBand / Spectrum-X)?

Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.

Can you help with NVIDIA AI Enterprise licensing?

Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.