Brand: NVIDIA | Category: GPUs
SKU: 920-23687-2570-000 | Part #: 920-23687-2570-000 | MPN: 920-23687-2570-000
Contact for Pricing — Request a Quote
The NVIDIA GB200 NVL72 is a rack-scale AI supercomputing system integrating 36 Grace CPUs and 72 Blackwell GPUs into a single NVLink-connected compute unit. Built on NVIDIA's fifth-generation NVLink fabric, the system delivers 1.4 exaflops of AI compute (FP4) and 720 petaflops of FP8 performance, enabling training and inference workloads at a scale previously achievable only by multi-rack HPC clusters. The 72 Blackwell B200 GPUs within the rack share a unified 8.6 TB aggregate HBM3e memory pool, eliminating traditional PCIe bottlenecks and allowing model parallelism across the entire rack as a single logical GPU domain.
The GB200 NVL72 is engineered for the most demanding large language model (LLM) training and inference workloads, including models exceeding one trillion parameters. Each B200 GPU incorporates the second-generation Transformer Engine with FP4 and FP8 mixed-precision support, dual-chip Blackwell GPU architecture per die, and fifth-generation NVLink providing 1.8 TB/s of bidirectional bandwidth per GPU. The 36 integrated Grace CPUs (Arm Neoverse V2 architecture) are connected to their paired B200 GPUs via NVLink-C2C, delivering 900 GB/s of coherent CPU-GPU interconnect bandwidth — eliminating the PCIe memory copy overhead that constrains conventional accelerated servers.
Designed for liquid-cooled data center deployments, the GB200 NVL72 utilizes a direct liquid cooling (DLC) architecture to manage the thermal demands of rack-scale AI compute. The system is well-suited for deployment in modern AI factories and hyperscale data centers equipped with rear-door or direct liquid cooling infrastructure. NVIDIA's NVLink Switch System within the rack enables all-to-all GPU communication at full bandwidth, making the GB200 NVL72 a foundational building block for sovereign AI infrastructure, cloud service provider GPU clusters, and enterprise-scale AI platforms across the UAE, GCC, EMEA, and APAC regions. Omnixon Global offers worldwide supply of this system to qualified enterprise, government, and research organizations.
| Manufacturer | NVIDIA |
| Manufacturer Part Number | 920-23687-2570-000 |
| Product Family | GB200 NVL72 |
| GPU Architecture | NVIDIA Blackwell |
| GPU Count per Rack | 72 × NVIDIA B200 GPUs |
| CPU Count per Rack | 36 × NVIDIA Grace CPUs (Arm Neoverse V2) |
| AI Performance (FP4) | 1.4 ExaFLOPS per rack |
| AI Performance (FP8) | 720 PFLOPS per rack |
| GPU Memory per GPU | 192 GB HBM3e |
| Total Aggregate GPU Memory | 8.6 TB HBM3e (across 72 GPUs) |
| GPU Memory Bandwidth per GPU | 8 TB/s |
| CPU-GPU Interconnect | NVLink-C2C at 900 GB/s bidirectional per Grace-Blackwell Superchip |
| NVLink GPU-to-GPU Bandwidth | 1.8 TB/s bidirectional per GPU |
| NVLink Generation | 5th Generation NVLink with NVLink Switch System |
| Interconnect Topology | NVLink Switch System enabling full all-to-all 72-GPU connectivity within the rack |
| Transformer Engine | 2nd Generation (FP4 and FP8 mixed precision support) |
| Cooling | Direct Liquid Cooling (DLC) |
| Form Factor | Rack-Scale System (full rack) |
| Target Infrastructure | Liquid-cooled AI data center / AI factory deployments |
| Supported Precision Formats | FP4, FP8, FP16, BF16, TF32, FP64 |
| Product Segment | Data Center / AI Infrastructure |
Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.
| Brand | NVIDIA |
| Category | GPUs |
| SKU | 920-23687-2570-000 |
| Part Number | 920-23687-2570-000 |
| Condition | New |
| Manufacturer Part Number | 920-23687-2570-000 |
| Product Family | GB200 NVL72 |
| GPU Architecture | NVIDIA Blackwell |
| GPU Count per Rack | 72 × NVIDIA B200 GPUs |
| CPU Count per Rack | 36 × NVIDIA Grace CPUs (Arm Neoverse V2) |
| AI Performance (FP4) | 1.4 ExaFLOPS per rack |
| AI Performance (FP8) | 720 PFLOPS per rack |
| GPU Memory per GPU | 192 GB HBM3e |
| Total Aggregate GPU Memory | 8.6 TB HBM3e (across 72 GPUs) |
| GPU Memory Bandwidth per GPU | 8 TB/s |
| CPU-GPU Interconnect | NVLink-C2C at 900 GB/s bidirectional per Grace-Blackwell Superchip |
| NVLink GPU-to-GPU Bandwidth | 1.8 TB/s bidirectional per GPU |
| NVLink Generation | 5th Generation NVLink with NVLink Switch System |
| Interconnect Topology | NVLink Switch System enabling full all-to-all 72-GPU connectivity within the rack |
| Transformer Engine | 2nd Generation (FP4 and FP8 mixed precision support) |
| Cooling | Direct Liquid Cooling (DLC) |
| Form Factor | Rack-Scale System (full rack) |
| Target Infrastructure | Liquid-cooled AI data center / AI factory deployments |
| Supported Precision Formats | FP4, FP8, FP16, BF16, TF32, FP64 |
| Product Segment | Data Center / AI Infrastructure |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.