Gigabyte NVIDIA GB200 NVL72 Rack-Scale GPU System

Gigabyte NVIDIA GB200 NVL72 Rack-Scale GPU System

Brand: Gigabyte | Category: GPUs

SKU: 900-21020-0040-000 | Part #: 900-21020-0040-000 | MPN: 900-21020-0040-000

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the Gigabyte NVIDIA GB200 NVL72 Rack-Scale GPU System

The Gigabyte NVIDIA GB200 NVL72 Rack-Scale GPU System is a full-rack AI supercomputing platform built around NVIDIA's Blackwell architecture, integrating 72 NVIDIA B200 Tensor Core GPUs and 36 NVIDIA Grace CPUs within a single NVLink-domain chassis. The system leverages fifth-generation NVLink to connect all 72 GPUs at an aggregate 130 TB/s of GPU-to-GPU bandwidth, enabling the entire rack to function as a single, unified compute entity rather than a collection of discrete nodes. With a total of 1.4 ExaFLOPS of AI performance (FP4) and 13.5 TB of unified HBM3e memory across the rack, the GB200 NVL72 is purpose-engineered for the most demanding large-scale model training and inference workloads in production datacenters.

At the hardware level, each B200 GPU in the system delivers up to 20 PetaFLOPS of FP4 Tensor Core performance and 192 GB of HBM3e memory operating at 8 TB/s of memory bandwidth per GPU. The Grace CPUs provide high-bandwidth, low-latency coherent memory interconnects via NVLink-C2C, eliminating traditional PCIe bottlenecks between CPU and GPU domains. The rack integrates a liquid-cooled thermal design capable of supporting the extreme power densities inherent to this class of system, with per-rack TDP in the range defined by NVIDIA's NVL72 reference platform. Networking is served by NVIDIA Quantum-X800 InfiniBand or Spectrum-X800 Ethernet fabric options for multi-rack scale-out.

The GB200 NVL72 targets hyperscale cloud operators, national AI research institutions, large financial services firms, and enterprise customers building sovereign AI infrastructure. It is optimized for trillion-parameter foundation model training, real-time large language model inference at scale, physics-based simulation, and high-throughput data analytics pipelines. Gigabyte's implementation adheres to the NVIDIA MGX reference architecture, ensuring ecosystem compatibility with NVIDIA AI Enterprise software, CUDA libraries, and the broader NVIDIA accelerated computing stack.

Ideal for

  • Large-scale foundation model training (LLM, multimodal, and diffusion models with trillions of parameters) requiring unified high-bandwidth GPU memory across the full rack
  • High-throughput generative AI inference serving, enabling real-time responses for enterprise-grade AI applications with minimal latency at massive concurrency
  • Sovereign AI datacenter buildout for government and national research bodies requiring on-premises, rack-scale accelerated compute capacity
  • Scientific high-performance computing including climate modeling, molecular dynamics simulation, and computational fluid dynamics at previously infeasible scales
  • Financial services risk modeling and quantitative analytics workloads demanding extreme floating-point throughput and large in-memory dataset processing
  • Healthcare and life sciences genomics, drug discovery, and medical imaging AI pipelines that benefit from unified memory capacity and NVLink-accelerated data movement

Technical specifications

ManufacturerGigabyte
Manufacturer Part Number900-21020-0040-000
System Form FactorFull rack (NVL72 rack-scale system)
GPU ArchitectureNVIDIA Blackwell
Number of GPUs72 x NVIDIA B200 Tensor Core GPU
Number of CPUs36 x NVIDIA Grace CPU
CPU-GPU InterconnectNVLink-C2C (coherent, high-bandwidth)
GPU Interconnect5th Generation NVLink
NVLink Aggregate Bandwidth130 TB/s (rack-wide GPU-to-GPU)
Total HBM3e Memory13.5 TB across 72 GPUs
HBM3e Memory per GPU192 GB
Memory Bandwidth per GPU8 TB/s
AI Performance (FP4, rack total)1.4 ExaFLOPS
AI Performance per GPU (FP4)20 PetaFLOPS
CoolingLiquid cooling (direct liquid cooling)
Networking OptionsNVIDIA Quantum-X800 InfiniBand / NVIDIA Spectrum-X800 Ethernet
Reference ArchitectureNVIDIA MGX
Software EcosystemNVIDIA AI Enterprise, CUDA, cuDNN, NCCL, TensorRT
Target DeploymentHyperscale datacenter, enterprise AI infrastructure, HPC
ComplianceNVIDIA NVL72 platform specification

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandGigabyte
CategoryGPUs
SKU900-21020-0040-000
Part Number900-21020-0040-000
ConditionNew
Manufacturer Part Number900-21020-0040-000
System Form FactorFull rack (NVL72 rack-scale system)
GPU ArchitectureNVIDIA Blackwell
Number of GPUs72 x NVIDIA B200 Tensor Core GPU
Number of CPUs36 x NVIDIA Grace CPU
CPU-GPU InterconnectNVLink-C2C (coherent, high-bandwidth)
GPU Interconnect5th Generation NVLink
NVLink Aggregate Bandwidth130 TB/s (rack-wide GPU-to-GPU)
Total HBM3e Memory13.5 TB across 72 GPUs
HBM3e Memory per GPU192 GB
Memory Bandwidth per GPU8 TB/s
AI Performance (FP4, rack total)1.4 ExaFLOPS
AI Performance per GPU (FP4)20 PetaFLOPS
CoolingLiquid cooling (direct liquid cooling)
Networking OptionsNVIDIA Quantum-X800 InfiniBand / NVIDIA Spectrum-X800 Ethernet
Reference ArchitectureNVIDIA MGX
Software EcosystemNVIDIA AI Enterprise, CUDA, cuDNN, NCCL, TensorRT
Target DeploymentHyperscale datacenter, enterprise AI infrastructure, HPC
ComplianceNVIDIA NVL72 platform specification

Frequently Asked Questions about Gigabyte NVIDIA GB200 NVL72 Rack-Scale GPU System

What server platforms accept the Gigabyte NVIDIA GB200 NVL72 Rack-Scale GPU System?

Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.

How long is the lead time on AI GPUs?

Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.

Do you supply matched networking (Quantum InfiniBand / Spectrum-X)?

Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.

Can you help with NVIDIA AI Enterprise licensing?

Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.