NVIDIA GB200 NVL72 Grace Blackwell Superchip

NVIDIA GB200 NVL72 Grace Blackwell Superchip

Brand: NVIDIA | Category: GPUs

SKU: NVID-93523B000000000 | Part #: 935-23B00-0000-000 | MPN: 935-23B00-0000-000

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the NVIDIA GB200 NVL72 Grace Blackwell Superchip

The NVIDIA GB200 NVL72 Grace Blackwell Superchip is a rack-scale AI computing platform that integrates 36 Grace Blackwell Superchips—each pairing one NVIDIA Grace CPU with two NVIDIA Blackwell B200 GPUs—into a single NVLink-connected rack system. The result is 72 Blackwell GPUs and 36 Grace CPUs operating as a unified, coherent compute fabric interconnected via fifth-generation NVLink at 1.8 TB/s of all-to-all GPU bandwidth within the rack. The platform delivers up to 1.4 exaflops of AI compute at FP4 precision, establishing a new class of infrastructure for the most demanding large-scale AI training and inference workloads.

The GB200 NVL72 leverages the Blackwell GPU architecture, which introduces a second-generation Transformer Engine, native FP4 and FP6 precision support, and a new Reliability, Availability, and Serviceability (RAS) Engine designed for continuous operation at scale. The Grace CPUs provide high-bandwidth, low-latency coherent memory access between CPU and GPU via NVLink-C2C at 900 GB/s per Superchip, eliminating PCIe bottlenecks. Each B200 GPU in the system incorporates 192 GB of HBM3e memory with 8 TB/s of memory bandwidth per GPU, enabling the platform to hold and serve extremely large model parameter sets entirely in high-bandwidth memory without offloading.

Designed explicitly for AI factory deployments, the GB200 NVL72 targets organizations running frontier large language model training, multi-trillion-parameter model inference, and national-scale AI infrastructure. The rack integrates liquid cooling as the standard thermal solution to manage the platform's power envelope in hyperscale and enterprise data center environments. Its NVLink Switch topology allows every GPU in the rack to communicate directly with every other GPU at full bandwidth, making it the preferred foundation for distributed deep learning jobs that cannot tolerate the latency or bandwidth constraints of traditional scale-out networking alone.

Ideal for

  • Training frontier large language models (LLMs) with hundreds of billions to trillions of parameters, where unified NVLink fabric eliminates inter-node bottlenecks within the rack
  • High-throughput inference serving of large generative AI models requiring low latency and massive aggregate memory capacity to hold full model weights on-device
  • Retrieval-augmented generation (RAG) pipelines and mixture-of-experts (MoE) model serving at enterprise and hyperscale production scale
  • Scientific simulation and digital twin workloads in climate modeling, drug discovery, and computational fluid dynamics that benefit from the platform's FP64 and mixed-precision capabilities
  • National AI infrastructure and sovereign AI factory deployments requiring maximum compute density and scalable multi-rack configurations via NVLink and InfiniBand fabrics
  • Continuous pre-training and fine-tuning pipelines for domain-specific foundation models in healthcare, financial services, and autonomous systems

Technical specifications

ManufacturerNVIDIA
Manufacturer Part Number935-23B00-0000-000
Platform NameNVIDIA GB200 NVL72
Form FactorRack-scale (full rack)
GPUs per Rack72 × NVIDIA Blackwell B200 GPUs
CPUs per Rack36 × NVIDIA Grace CPUs (ARM Neoverse V2-based)
GPU ArchitectureNVIDIA Blackwell
GPU Memory per B200192 GB HBM3e
Total GPU Memory (Rack)13.5 TB HBM3e
GPU Memory Bandwidth per B2008 TB/s
Total GPU Memory Bandwidth (Rack)576 TB/s
AI Compute (FP4, Rack)1.4 exaflops
AI Compute (FP8, Rack)720 petaflops
CPU-GPU InterconnectNVLink-C2C at 900 GB/s per Grace Blackwell Superchip (bidirectional)
Intra-Rack GPU InterconnectNVIDIA NVLink 5 (fifth generation), 1.8 TB/s all-to-all aggregate bandwidth
NVLink Switch GenerationNVLink Switch (fifth generation)
External Fabric SupportNVIDIA InfiniBand and Ethernet (for multi-rack scale-out)
Thermal SolutionLiquid cooling (required)
Transformer Engine GenerationSecond generation (native FP4, FP6, FP8, BF16, FP16, FP32, TF32)

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandNVIDIA
CategoryGPUs
SKUNVID-93523B000000000
Part Number935-23B00-0000-000
ConditionNew
Manufacturer Part Number935-23B00-0000-000
Platform NameNVIDIA GB200 NVL72
Form FactorRack-scale (full rack)
GPUs per Rack72 × NVIDIA Blackwell B200 GPUs
CPUs per Rack36 × NVIDIA Grace CPUs (ARM Neoverse V2-based)
GPU ArchitectureNVIDIA Blackwell
GPU Memory per B200192 GB HBM3e
Total GPU Memory (Rack)13.5 TB HBM3e
GPU Memory Bandwidth per B2008 TB/s
Total GPU Memory Bandwidth (Rack)576 TB/s
AI Compute (FP4, Rack)1.4 exaflops
AI Compute (FP8, Rack)720 petaflops
CPU-GPU InterconnectNVLink-C2C at 900 GB/s per Grace Blackwell Superchip (bidirectional)
Intra-Rack GPU InterconnectNVIDIA NVLink 5 (fifth generation), 1.8 TB/s all-to-all aggregate bandwidth
NVLink Switch GenerationNVLink Switch (fifth generation)
External Fabric SupportNVIDIA InfiniBand and Ethernet (for multi-rack scale-out)
Thermal SolutionLiquid cooling (required)
Transformer Engine GenerationSecond generation (native FP4, FP6, FP8, BF16, FP16, FP32, TF32)