NVIDIA GB200 NVL36 Rack-Scale System 6912GB HBM3e

NVIDIA GB200 NVL36 Rack-Scale System 6912GB HBM3e

Brand: NVIDIA | Category: GPUs

SKU: GB200-NVL36 | Part #: GB200-NVL36 | MPN: GB200-NVL36

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the NVIDIA GB200 NVL36 Rack-Scale System 6912GB HBM3e

The NVIDIA GB200 NVL36 is a rack-scale AI computing system integrating 36 Blackwell GPUs and 18 Grace CPUs within a single NVLink-connected rack unit, delivering unified high-bandwidth memory and compute resources across the entire chassis. Built on the NVIDIA Blackwell GPU architecture, each GB200 Superchip pairs two B200 GPUs with one Grace CPU via NVLink-C2C interconnect, achieving up to 900 GB/s chip-to-chip bandwidth. The rack system aggregates 6,912 GB of HBM3e memory across all GPUs, providing an unprecedented pooled memory capacity for training and inferencing extremely large AI models that previously required multi-rack deployments.

The GB200 NVL36 delivers up to 1.4 exaflops of AI compute (FP4) across the full rack, with NVLink Switch System fabric enabling all 36 GPUs to communicate at 1.8 TB/s of all-to-all GPU bandwidth. This tight integration eliminates traditional PCIe bottlenecks and allows the rack to operate as a single logical unit for distributed AI workloads. The system is purpose-built for the most demanding large language model (LLM) training runs, mixture-of-experts (MoE) architectures, and real-time inference serving at scale, where inter-GPU communication latency and memory capacity are the primary constraints.

Designed for modern AI-first data centers, the GB200 NVL36 uses a liquid-cooled architecture to manage the substantial thermal output of a full rack of Blackwell Superchips, enabling high-density deployment within standard data center footprints. The system integrates with NVIDIA's full software stack including CUDA, cuDNN, TensorRT, and NEMO frameworks, and supports NVIDIA Confidential Computing for enterprise security requirements. Organizations across cloud service providers, enterprise AI research, financial services, and life sciences rely on this platform to consolidate AI infrastructure and reduce the total number of nodes required for frontier model workloads.

Ideal for

  • Training frontier-scale large language models (LLMs) with hundreds of billions to trillions of parameters, leveraging the 6,912 GB pooled HBM3e memory to keep model weights and optimizer states on-chip across the full rack
  • High-throughput LLM inference serving for enterprise AI applications, using the NVLink-connected GPU pool to minimize token generation latency at scale for internal or customer-facing deployments
  • Mixture-of-experts (MoE) model training and inference where large aggregate memory and ultra-low-latency all-to-all communication between GPUs are critical to maintaining expert routing efficiency
  • Scientific simulation and HPC workloads in pharmaceutical drug discovery, climate modeling, and genomics that require tightly coupled GPU compute with large shared memory spaces
  • Enterprise AI model fine-tuning pipelines for domain-specific foundation models in financial services, healthcare, and manufacturing, where rapid iteration over large proprietary datasets is required
  • Sovereign AI and national AI infrastructure deployments requiring dense, rack-scale compute capacity with support for confidential computing and enterprise-grade data governance

Technical specifications

ManufacturerNVIDIA
Manufacturer Part NumberGB200-NVL36
GPU ArchitectureNVIDIA Blackwell
System Configuration36 × B200 GPUs + 18 × Grace CPUs (18 × GB200 Superchips)
Total HBM3e Memory6,912 GB
HBM3e Memory per GPU192 GB
HBM3e Memory Bandwidth per GPU8 TB/s
Total Memory Bandwidth (Rack)Up to 288 TB/s aggregate
Peak AI Compute (FP4, Rack)Up to 1.4 Exaflops
Peak AI Compute per GPU (FP4)Up to 40 PFLOPS
NVLink All-to-All GPU Bandwidth1.8 TB/s (bidirectional) across all 36 GPUs
NVLink-C2C Bandwidth (Grace–Blackwell)900 GB/s bidirectional per Superchip
Interconnect FabricNVLink Switch System (fifth generation NVLink)
Grace CPU ArchitectureArm Neoverse V2, 72 cores per CPU
Grace CPU MemoryLPDDR5X with up to 480 GB/s per CPU
CoolingLiquid cooling (direct liquid cooling required)
Form FactorRack-scale system (full rack)
Supported Precision FormatsFP4, FP8, FP16, BF16, TF32, FP32, INT8
Confidential ComputingSupported (NVIDIA Hopper Confidential Computing architecture extended in Blackwell)
Software StackCUDA, cuDNN, TensorRT, NEMO, Triton Inference Server, NVIDIA AI Enterprise

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandNVIDIA
CategoryGPUs
SKUGB200-NVL36
Part NumberGB200-NVL36
ConditionNew
Manufacturer Part NumberGB200-NVL36
GPU ArchitectureNVIDIA Blackwell
System Configuration36 × B200 GPUs + 18 × Grace CPUs (18 × GB200 Superchips)
Total HBM3e Memory6,912 GB
HBM3e Memory per GPU192 GB
HBM3e Memory Bandwidth per GPU8 TB/s
Total Memory Bandwidth (Rack)Up to 288 TB/s aggregate
Peak AI Compute (FP4, Rack)Up to 1.4 Exaflops
Peak AI Compute per GPU (FP4)Up to 40 PFLOPS
NVLink All-to-All GPU Bandwidth1.8 TB/s (bidirectional) across all 36 GPUs
NVLink-C2C Bandwidth (Grace–Blackwell)900 GB/s bidirectional per Superchip
Interconnect FabricNVLink Switch System (fifth generation NVLink)
Grace CPU ArchitectureArm Neoverse V2, 72 cores per CPU
Grace CPU MemoryLPDDR5X with up to 480 GB/s per CPU
CoolingLiquid cooling (direct liquid cooling required)
Form FactorRack-scale system (full rack)
Supported Precision FormatsFP4, FP8, FP16, BF16, TF32, FP32, INT8
Confidential ComputingSupported (NVIDIA Hopper Confidential Computing architecture extended in Blackwell)
Software StackCUDA, cuDNN, TensorRT, NEMO, Triton Inference Server, NVIDIA AI Enterprise

Frequently Asked Questions about NVIDIA GB200 NVL36 Rack-Scale System 6912GB HBM3e

What server platforms accept the NVIDIA GB200 NVL36 Rack-Scale System 6912GB HBM3e?

Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.

How long is the lead time on AI GPUs?

Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.

Do you supply matched networking (Quantum InfiniBand / Spectrum-X)?

Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.

Can you help with NVIDIA AI Enterprise licensing?

Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.