Brand: Gigabyte | Category: GPUs
SKU: GV-NGB200NVL72 | Part #: GV-NGB200NVL72 | MPN: GV-NGB200NVL72
Contact for Pricing — Request a Quote
The GIGABYTE GB200 NVL72 Grace Blackwell Superchip Rack Solution (GV-NGB200NVL72) is a full-rack AI computing platform built around NVIDIA's GB200 NVL72 architecture, integrating 36 Grace Blackwell Superchips — each pairing one NVIDIA Grace CPU with two NVIDIA Blackwell B200 GPUs — for a total of 72 Blackwell GPUs and 36 Grace Arm-based CPUs within a single NVLink-interconnected rack. The system leverages NVLink Switch fabric to unify all 72 GPUs into a single, coherent high-bandwidth memory and compute domain, delivering up to 1.4 exaFLOPS of AI compute at FP4 precision across the rack. This tightly integrated design eliminates traditional PCIe bottlenecks by enabling GPU-to-GPU communication at NVLink speeds, making it purpose-built for the most demanding large-scale model training and inference workloads in modern datacenters.
Each Grace Blackwell Superchip within the rack combines the Grace CPU's 72 Arm Neoverse V2 cores with the high-bandwidth LPDDR5X CPU memory subsystem and the Blackwell GPU architecture's second-generation Transformer Engine, supporting FP4, FP8, FP16, BF16, TF32, and FP64 precision modes. The rack solution incorporates NVIDIA's liquid-cooled thermal management design, supporting direct liquid cooling (DLC) as the primary thermal pathway to sustain the extreme power envelope of the full NVL72 configuration. GIGABYTE's integration of this platform is designed to meet stringent datacenter power and mechanical infrastructure requirements, with the rack solution fitting within standard 19-inch datacenter rack form factors under defined structural and coolant supply specifications.
Targeted at hyperscale cloud operators, national AI research institutions, and large enterprise AI infrastructure deployments, the GB200 NVL72 rack solution from GIGABYTE is positioned for workloads where GPU memory capacity, inter-GPU bandwidth, and sustained compute throughput at scale are the primary constraints. The NVL72 configuration provides a combined GPU high-bandwidth memory capacity of 13.5 TB across all 72 B200 GPUs, enabling the training and inference of frontier-scale large language models, multimodal foundation models, and scientific simulation workloads that exceed the memory capacity of any single-node or smaller multi-GPU system. GIGABYTE's global datacenter infrastructure expertise ensures this rack solution is validated for deployment in enterprise and hyperscale datacenter environments across regions including UAE, GCC, EMEA, and APAC.
| Manufacturer | Gigabyte |
| Manufacturer Part Number | GV-NGB200NVL72 |
| Platform Architecture | NVIDIA GB200 NVL72 Grace Blackwell Superchip |
| Total GPUs per Rack | 72 × NVIDIA Blackwell B200 GPUs |
| Total CPUs per Rack | 36 × NVIDIA Grace CPUs (72 Arm Neoverse V2 cores each) |
| Superchip Configuration | 36 × GB200 Grace Blackwell Superchips (1 Grace CPU + 2 B200 GPUs per Superchip) |
| GPU Interconnect Fabric | NVLink Switch (5th generation NVLink), full NVL72 all-to-all GPU fabric |
| NVLink Bandwidth (Rack Total) | Up to 130 TB/s bisection bandwidth across the NVLink switch fabric |
| Peak AI Compute (FP4) | Up to 1.4 ExaFLOPS per rack |
| Peak AI Compute (FP8) | Up to 720 PFLOPS per rack |
| GPU Memory Type | HBM3e |
| Total GPU HBM Capacity (Rack) | 13.5 TB (72 × 192 GB HBM3e per B200 GPU) |
| GPU Memory Bandwidth (per B200) | 8 TB/s HBM3e memory bandwidth |
| CPU Memory Type | LPDDR5X with NVLink-C2C CPU-to-GPU interconnect |
| Supported Precision Formats | FP4, FP8, FP16, BF16, TF32, FP64, INT8 |
| Cooling Method | Direct Liquid Cooling (DLC) — primary thermal pathway |
| Form Factor | Full rack, 19-inch datacenter rack compatible |
| GPU Architecture Generation | NVIDIA Blackwell (second-generation Transformer Engine) |
| Primary Software Ecosystem | NVIDIA CUDA, NCCL, cuDNN, TensorRT, NIM, NEMO, Triton Inference Server |
| Target Deployment Environment | Hyperscale datacenter, enterprise AI infrastructure, HPC centers |
Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.
| Brand | Gigabyte |
| Category | GPUs |
| SKU | GV-NGB200NVL72 |
| Part Number | GV-NGB200NVL72 |
| Condition | New |
| Manufacturer Part Number | GV-NGB200NVL72 |
| Platform Architecture | NVIDIA GB200 NVL72 Grace Blackwell Superchip |
| Total GPUs per Rack | 72 × NVIDIA Blackwell B200 GPUs |
| Total CPUs per Rack | 36 × NVIDIA Grace CPUs (72 Arm Neoverse V2 cores each) |
| Superchip Configuration | 36 × GB200 Grace Blackwell Superchips (1 Grace CPU + 2 B200 GPUs per Superchip) |
| GPU Interconnect Fabric | NVLink Switch (5th generation NVLink), full NVL72 all-to-all GPU fabric |
| NVLink Bandwidth (Rack Total) | Up to 130 TB/s bisection bandwidth across the NVLink switch fabric |
| Peak AI Compute (FP4) | Up to 1.4 ExaFLOPS per rack |
| Peak AI Compute (FP8) | Up to 720 PFLOPS per rack |
| GPU Memory Type | HBM3e |
| Total GPU HBM Capacity (Rack) | 13.5 TB (72 × 192 GB HBM3e per B200 GPU) |
| GPU Memory Bandwidth (per B200) | 8 TB/s HBM3e memory bandwidth |
| CPU Memory Type | LPDDR5X with NVLink-C2C CPU-to-GPU interconnect |
| Supported Precision Formats | FP4, FP8, FP16, BF16, TF32, FP64, INT8 |
| Cooling Method | Direct Liquid Cooling (DLC) — primary thermal pathway |
| Form Factor | Full rack, 19-inch datacenter rack compatible |
| GPU Architecture Generation | NVIDIA Blackwell (second-generation Transformer Engine) |
| Primary Software Ecosystem | NVIDIA CUDA, NCCL, cuDNN, TensorRT, NIM, NEMO, Triton Inference Server |
| Target Deployment Environment | Hyperscale datacenter, enterprise AI infrastructure, HPC centers |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.