Brand: Gigabyte | Category: GPUs
SKU: 900-21020-0040-000 | Part #: 900-21020-0040-000 | MPN: 900-21020-0040-000
Contact for Pricing — Request a Quote
The Gigabyte NVIDIA GB200 NVL72 Rack-Scale GPU System is a full-rack AI supercomputing platform built around NVIDIA's Blackwell architecture, integrating 72 NVIDIA B200 Tensor Core GPUs and 36 NVIDIA Grace CPUs within a single NVLink-domain chassis. The system leverages fifth-generation NVLink to connect all 72 GPUs at an aggregate 130 TB/s of GPU-to-GPU bandwidth, enabling the entire rack to function as a single, unified compute entity rather than a collection of discrete nodes. With a total of 1.4 ExaFLOPS of AI performance (FP4) and 13.5 TB of unified HBM3e memory across the rack, the GB200 NVL72 is purpose-engineered for the most demanding large-scale model training and inference workloads in production datacenters.
At the hardware level, each B200 GPU in the system delivers up to 20 PetaFLOPS of FP4 Tensor Core performance and 192 GB of HBM3e memory operating at 8 TB/s of memory bandwidth per GPU. The Grace CPUs provide high-bandwidth, low-latency coherent memory interconnects via NVLink-C2C, eliminating traditional PCIe bottlenecks between CPU and GPU domains. The rack integrates a liquid-cooled thermal design capable of supporting the extreme power densities inherent to this class of system, with per-rack TDP in the range defined by NVIDIA's NVL72 reference platform. Networking is served by NVIDIA Quantum-X800 InfiniBand or Spectrum-X800 Ethernet fabric options for multi-rack scale-out.
The GB200 NVL72 targets hyperscale cloud operators, national AI research institutions, large financial services firms, and enterprise customers building sovereign AI infrastructure. It is optimized for trillion-parameter foundation model training, real-time large language model inference at scale, physics-based simulation, and high-throughput data analytics pipelines. Gigabyte's implementation adheres to the NVIDIA MGX reference architecture, ensuring ecosystem compatibility with NVIDIA AI Enterprise software, CUDA libraries, and the broader NVIDIA accelerated computing stack.
| Manufacturer | Gigabyte |
| Manufacturer Part Number | 900-21020-0040-000 |
| System Form Factor | Full rack (NVL72 rack-scale system) |
| GPU Architecture | NVIDIA Blackwell |
| Number of GPUs | 72 x NVIDIA B200 Tensor Core GPU |
| Number of CPUs | 36 x NVIDIA Grace CPU |
| CPU-GPU Interconnect | NVLink-C2C (coherent, high-bandwidth) |
| GPU Interconnect | 5th Generation NVLink |
| NVLink Aggregate Bandwidth | 130 TB/s (rack-wide GPU-to-GPU) |
| Total HBM3e Memory | 13.5 TB across 72 GPUs |
| HBM3e Memory per GPU | 192 GB |
| Memory Bandwidth per GPU | 8 TB/s |
| AI Performance (FP4, rack total) | 1.4 ExaFLOPS |
| AI Performance per GPU (FP4) | 20 PetaFLOPS |
| Cooling | Liquid cooling (direct liquid cooling) |
| Networking Options | NVIDIA Quantum-X800 InfiniBand / NVIDIA Spectrum-X800 Ethernet |
| Reference Architecture | NVIDIA MGX |
| Software Ecosystem | NVIDIA AI Enterprise, CUDA, cuDNN, NCCL, TensorRT |
| Target Deployment | Hyperscale datacenter, enterprise AI infrastructure, HPC |
| Compliance | NVIDIA NVL72 platform specification |
Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.
| Brand | Gigabyte |
| Category | GPUs |
| SKU | 900-21020-0040-000 |
| Part Number | 900-21020-0040-000 |
| Condition | New |
| Manufacturer Part Number | 900-21020-0040-000 |
| System Form Factor | Full rack (NVL72 rack-scale system) |
| GPU Architecture | NVIDIA Blackwell |
| Number of GPUs | 72 x NVIDIA B200 Tensor Core GPU |
| Number of CPUs | 36 x NVIDIA Grace CPU |
| CPU-GPU Interconnect | NVLink-C2C (coherent, high-bandwidth) |
| GPU Interconnect | 5th Generation NVLink |
| NVLink Aggregate Bandwidth | 130 TB/s (rack-wide GPU-to-GPU) |
| Total HBM3e Memory | 13.5 TB across 72 GPUs |
| HBM3e Memory per GPU | 192 GB |
| Memory Bandwidth per GPU | 8 TB/s |
| AI Performance (FP4, rack total) | 1.4 ExaFLOPS |
| AI Performance per GPU (FP4) | 20 PetaFLOPS |
| Cooling | Liquid cooling (direct liquid cooling) |
| Networking Options | NVIDIA Quantum-X800 InfiniBand / NVIDIA Spectrum-X800 Ethernet |
| Reference Architecture | NVIDIA MGX |
| Software Ecosystem | NVIDIA AI Enterprise, CUDA, cuDNN, NCCL, TensorRT |
| Target Deployment | Hyperscale datacenter, enterprise AI infrastructure, HPC |
| Compliance | NVIDIA NVL72 platform specification |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.