Brand: NVIDIA | Category: GPUs
SKU: 900-21761-0040-000 | Part #: 900-21761-0040-000 | MPN: 900-21761-0040-000
Contact for Pricing — Request a Quote
The NVIDIA GB200 NVL72 is a rack-scale AI supercomputing system built on the Blackwell Ultra architecture, integrating 36 GB200 Superchips into a single liquid-cooled rack unit. Each GB200 Superchip pairs two NVIDIA B200 Tensor Core GPUs with one NVIDIA Grace CPU via NVLink-C2C, delivering a tightly unified compute fabric across the entire rack. The system interconnects all 72 B200 GPUs using fifth-generation NVLink switching, creating a single logical GPU with 1.4 TB of aggregate HBM3e memory and 130 TB/s of all-to-all GPU bandwidth — enabling model parallelism at a scale previously requiring multi-rack configurations.
Designed specifically for the demands of large language model training and inference, the GB200 NVL72 delivers up to 1.4 ExaFLOPS of AI performance at FP4 precision and 720 PFLOPS at FP8, with full support for FP16, BF16, TF32, and FP64 compute modes. The Grace CPUs within the system provide 144 high-performance Arm Neoverse V2 cores and up to 17.6 TB/s of CPU-to-GPU bandwidth via NVLink-C2C, eliminating traditional PCIe bottlenecks and enabling direct, high-bandwidth access to GPU memory from the host processing layer. The system ships as a complete rack-scale reference design inclusive of NVLink switches, high-speed networking interfaces, power infrastructure, and direct liquid cooling (DLC) subsystems.
The GB200 NVL72 is purpose-built for hyperscale data centers and enterprise AI infrastructure teams deploying frontier-scale workloads including trillion-parameter model training, real-time inference serving at massive throughput, and advanced scientific simulation. It is the foundational compute platform for customers building sovereign AI infrastructure, next-generation cloud services, and high-performance computing (HPC) clusters demanding maximum GPU-to-GPU communication bandwidth in a rack-contained form factor.
| Manufacturer | NVIDIA |
| Manufacturer Part Number | 900-21761-0040-000 |
| Product Family | Blackwell |
| GPU Architecture | NVIDIA Blackwell |
| System Configuration | 72x NVIDIA B200 Tensor Core GPUs, 36x NVIDIA Grace CPUs (36 GB200 Superchips) |
| Total GPU Memory | 1.4 TB HBM3e |
| GPU Memory Per B200 | 192 GB HBM3e |
| GPU Memory Bandwidth Per B200 | 8 TB/s |
| Total AI Performance (FP4) | 1.4 ExaFLOPS |
| Total AI Performance (FP8) | 720 PFLOPS |
| GPU Interconnect | NVLink 5 (fifth-generation), 130 TB/s all-to-all bisection bandwidth |
| CPU-to-GPU Interconnect | NVLink-C2C, up to 900 GB/s per direction per Superchip |
| CPU | NVIDIA Grace (Arm Neoverse V2), 72 cores per CPU, 144 CPU cores total per Superchip pair |
| CPU Memory | LPDDR5X per Grace CPU |
| Supported Precisions | FP4, FP8, FP16, BF16, TF32, FP32, FP64, INT8 |
| Cooling | Direct Liquid Cooling (DLC) |
| Form Factor | Rack-scale system (full rack) |
| Networking | NVIDIA ConnectX-7 / BlueField-3 high-speed networking interfaces |
| Operating System Support | Linux (Ubuntu, RHEL, and compatible enterprise distributions) |
| CUDA Compatibility | CUDA 12.x and later |
Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.
| Brand | NVIDIA |
| Category | GPUs |
| SKU | 900-21761-0040-000 |
| Part Number | 900-21761-0040-000 |
| Condition | New |
| Manufacturer Part Number | 900-21761-0040-000 |
| Product Family | Blackwell |
| GPU Architecture | NVIDIA Blackwell |
| System Configuration | 72x NVIDIA B200 Tensor Core GPUs, 36x NVIDIA Grace CPUs (36 GB200 Superchips) |
| Total GPU Memory | 1.4 TB HBM3e |
| GPU Memory Per B200 | 192 GB HBM3e |
| GPU Memory Bandwidth Per B200 | 8 TB/s |
| Total AI Performance (FP4) | 1.4 ExaFLOPS |
| Total AI Performance (FP8) | 720 PFLOPS |
| GPU Interconnect | NVLink 5 (fifth-generation), 130 TB/s all-to-all bisection bandwidth |
| CPU-to-GPU Interconnect | NVLink-C2C, up to 900 GB/s per direction per Superchip |
| CPU | NVIDIA Grace (Arm Neoverse V2), 72 cores per CPU, 144 CPU cores total per Superchip pair |
| CPU Memory | LPDDR5X per Grace CPU |
| Supported Precisions | FP4, FP8, FP16, BF16, TF32, FP32, FP64, INT8 |
| Cooling | Direct Liquid Cooling (DLC) |
| Form Factor | Rack-scale system (full rack) |
| Networking | NVIDIA ConnectX-7 / BlueField-3 high-speed networking interfaces |
| Operating System Support | Linux (Ubuntu, RHEL, and compatible enterprise distributions) |
| CUDA Compatibility | CUDA 12.x and later |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.