GIGABYTE GB200 NVL72 Grace Blackwell Superchip Rack Solution

GIGABYTE GB200 NVL72 Grace Blackwell Superchip Rack Solution

Brand: Gigabyte | Category: GPUs

SKU: GV-NGB200NVL72 | Part #: GV-NGB200NVL72 | MPN: GV-NGB200NVL72

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the GIGABYTE GB200 NVL72 Grace Blackwell Superchip Rack Solution

The GIGABYTE GB200 NVL72 Grace Blackwell Superchip Rack Solution (GV-NGB200NVL72) is a full-rack AI computing platform built around NVIDIA's GB200 NVL72 architecture, integrating 36 Grace Blackwell Superchips — each pairing one NVIDIA Grace CPU with two NVIDIA Blackwell B200 GPUs — for a total of 72 Blackwell GPUs and 36 Grace Arm-based CPUs within a single NVLink-interconnected rack. The system leverages NVLink Switch fabric to unify all 72 GPUs into a single, coherent high-bandwidth memory and compute domain, delivering up to 1.4 exaFLOPS of AI compute at FP4 precision across the rack. This tightly integrated design eliminates traditional PCIe bottlenecks by enabling GPU-to-GPU communication at NVLink speeds, making it purpose-built for the most demanding large-scale model training and inference workloads in modern datacenters.

Each Grace Blackwell Superchip within the rack combines the Grace CPU's 72 Arm Neoverse V2 cores with the high-bandwidth LPDDR5X CPU memory subsystem and the Blackwell GPU architecture's second-generation Transformer Engine, supporting FP4, FP8, FP16, BF16, TF32, and FP64 precision modes. The rack solution incorporates NVIDIA's liquid-cooled thermal management design, supporting direct liquid cooling (DLC) as the primary thermal pathway to sustain the extreme power envelope of the full NVL72 configuration. GIGABYTE's integration of this platform is designed to meet stringent datacenter power and mechanical infrastructure requirements, with the rack solution fitting within standard 19-inch datacenter rack form factors under defined structural and coolant supply specifications.

Targeted at hyperscale cloud operators, national AI research institutions, and large enterprise AI infrastructure deployments, the GB200 NVL72 rack solution from GIGABYTE is positioned for workloads where GPU memory capacity, inter-GPU bandwidth, and sustained compute throughput at scale are the primary constraints. The NVL72 configuration provides a combined GPU high-bandwidth memory capacity of 13.5 TB across all 72 B200 GPUs, enabling the training and inference of frontier-scale large language models, multimodal foundation models, and scientific simulation workloads that exceed the memory capacity of any single-node or smaller multi-GPU system. GIGABYTE's global datacenter infrastructure expertise ensures this rack solution is validated for deployment in enterprise and hyperscale datacenter environments across regions including UAE, GCC, EMEA, and APAC.

Ideal for

  • Training and fine-tuning frontier-scale large language models (LLMs) with hundreds of billions to trillions of parameters using the unified 13.5 TB NVLink-pooled GPU memory domain
  • High-throughput generative AI inference serving for enterprise-scale deployments requiring maximum tokens-per-second output with minimal latency across concurrent requests
  • Large-scale multimodal foundation model development combining vision, language, and structured data modalities that demand massive parallel compute and high inter-GPU bandwidth
  • Scientific and HPC simulation workloads in climate modeling, genomics, drug discovery, and computational fluid dynamics that benefit from FP64 precision at extreme GPU core counts
  • National and institutional AI research infrastructure requiring a single-rack, fully interconnected GPU fabric that operates as a unified compute node for distributed training frameworks
  • Sovereign AI datacenter buildouts in GCC and EMEA regions requiring validated, enterprise-grade rack-scale GPU infrastructure with liquid cooling support for high-density facilities

Technical specifications

ManufacturerGigabyte
Manufacturer Part NumberGV-NGB200NVL72
Platform ArchitectureNVIDIA GB200 NVL72 Grace Blackwell Superchip
Total GPUs per Rack72 × NVIDIA Blackwell B200 GPUs
Total CPUs per Rack36 × NVIDIA Grace CPUs (72 Arm Neoverse V2 cores each)
Superchip Configuration36 × GB200 Grace Blackwell Superchips (1 Grace CPU + 2 B200 GPUs per Superchip)
GPU Interconnect FabricNVLink Switch (5th generation NVLink), full NVL72 all-to-all GPU fabric
NVLink Bandwidth (Rack Total)Up to 130 TB/s bisection bandwidth across the NVLink switch fabric
Peak AI Compute (FP4)Up to 1.4 ExaFLOPS per rack
Peak AI Compute (FP8)Up to 720 PFLOPS per rack
GPU Memory TypeHBM3e
Total GPU HBM Capacity (Rack)13.5 TB (72 × 192 GB HBM3e per B200 GPU)
GPU Memory Bandwidth (per B200)8 TB/s HBM3e memory bandwidth
CPU Memory TypeLPDDR5X with NVLink-C2C CPU-to-GPU interconnect
Supported Precision FormatsFP4, FP8, FP16, BF16, TF32, FP64, INT8
Cooling MethodDirect Liquid Cooling (DLC) — primary thermal pathway
Form FactorFull rack, 19-inch datacenter rack compatible
GPU Architecture GenerationNVIDIA Blackwell (second-generation Transformer Engine)
Primary Software EcosystemNVIDIA CUDA, NCCL, cuDNN, TensorRT, NIM, NEMO, Triton Inference Server
Target Deployment EnvironmentHyperscale datacenter, enterprise AI infrastructure, HPC centers

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandGigabyte
CategoryGPUs
SKUGV-NGB200NVL72
Part NumberGV-NGB200NVL72
ConditionNew
Manufacturer Part NumberGV-NGB200NVL72
Platform ArchitectureNVIDIA GB200 NVL72 Grace Blackwell Superchip
Total GPUs per Rack72 × NVIDIA Blackwell B200 GPUs
Total CPUs per Rack36 × NVIDIA Grace CPUs (72 Arm Neoverse V2 cores each)
Superchip Configuration36 × GB200 Grace Blackwell Superchips (1 Grace CPU + 2 B200 GPUs per Superchip)
GPU Interconnect FabricNVLink Switch (5th generation NVLink), full NVL72 all-to-all GPU fabric
NVLink Bandwidth (Rack Total)Up to 130 TB/s bisection bandwidth across the NVLink switch fabric
Peak AI Compute (FP4)Up to 1.4 ExaFLOPS per rack
Peak AI Compute (FP8)Up to 720 PFLOPS per rack
GPU Memory TypeHBM3e
Total GPU HBM Capacity (Rack)13.5 TB (72 × 192 GB HBM3e per B200 GPU)
GPU Memory Bandwidth (per B200)8 TB/s HBM3e memory bandwidth
CPU Memory TypeLPDDR5X with NVLink-C2C CPU-to-GPU interconnect
Supported Precision FormatsFP4, FP8, FP16, BF16, TF32, FP64, INT8
Cooling MethodDirect Liquid Cooling (DLC) — primary thermal pathway
Form FactorFull rack, 19-inch datacenter rack compatible
GPU Architecture GenerationNVIDIA Blackwell (second-generation Transformer Engine)
Primary Software EcosystemNVIDIA CUDA, NCCL, cuDNN, TensorRT, NIM, NEMO, Triton Inference Server
Target Deployment EnvironmentHyperscale datacenter, enterprise AI infrastructure, HPC centers

Frequently Asked Questions about GIGABYTE GB200 NVL72 Grace Blackwell Superchip Rack Solution

What server platforms accept the GIGABYTE GB200 NVL72 Grace Blackwell Superchip Rack Solution?

Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.

How long is the lead time on AI GPUs?

Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.

Do you supply matched networking (Quantum InfiniBand / Spectrum-X)?

Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.

Can you help with NVIDIA AI Enterprise licensing?

Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.