Brand: NVIDIA | Category: GPUs
SKU: GB200-NVL36 | Part #: GB200-NVL36 | MPN: GB200-NVL36
Contact for Pricing — Request a Quote
The NVIDIA GB200 NVL36 is a rack-scale AI computing system integrating 36 Blackwell GPUs and 18 Grace CPUs within a single NVLink-connected rack unit, delivering unified high-bandwidth memory and compute resources across the entire chassis. Built on the NVIDIA Blackwell GPU architecture, each GB200 Superchip pairs two B200 GPUs with one Grace CPU via NVLink-C2C interconnect, achieving up to 900 GB/s chip-to-chip bandwidth. The rack system aggregates 6,912 GB of HBM3e memory across all GPUs, providing an unprecedented pooled memory capacity for training and inferencing extremely large AI models that previously required multi-rack deployments.
The GB200 NVL36 delivers up to 1.4 exaflops of AI compute (FP4) across the full rack, with NVLink Switch System fabric enabling all 36 GPUs to communicate at 1.8 TB/s of all-to-all GPU bandwidth. This tight integration eliminates traditional PCIe bottlenecks and allows the rack to operate as a single logical unit for distributed AI workloads. The system is purpose-built for the most demanding large language model (LLM) training runs, mixture-of-experts (MoE) architectures, and real-time inference serving at scale, where inter-GPU communication latency and memory capacity are the primary constraints.
Designed for modern AI-first data centers, the GB200 NVL36 uses a liquid-cooled architecture to manage the substantial thermal output of a full rack of Blackwell Superchips, enabling high-density deployment within standard data center footprints. The system integrates with NVIDIA's full software stack including CUDA, cuDNN, TensorRT, and NEMO frameworks, and supports NVIDIA Confidential Computing for enterprise security requirements. Organizations across cloud service providers, enterprise AI research, financial services, and life sciences rely on this platform to consolidate AI infrastructure and reduce the total number of nodes required for frontier model workloads.
| Manufacturer | NVIDIA |
| Manufacturer Part Number | GB200-NVL36 |
| GPU Architecture | NVIDIA Blackwell |
| System Configuration | 36 × B200 GPUs + 18 × Grace CPUs (18 × GB200 Superchips) |
| Total HBM3e Memory | 6,912 GB |
| HBM3e Memory per GPU | 192 GB |
| HBM3e Memory Bandwidth per GPU | 8 TB/s |
| Total Memory Bandwidth (Rack) | Up to 288 TB/s aggregate |
| Peak AI Compute (FP4, Rack) | Up to 1.4 Exaflops |
| Peak AI Compute per GPU (FP4) | Up to 40 PFLOPS |
| NVLink All-to-All GPU Bandwidth | 1.8 TB/s (bidirectional) across all 36 GPUs |
| NVLink-C2C Bandwidth (Grace–Blackwell) | 900 GB/s bidirectional per Superchip |
| Interconnect Fabric | NVLink Switch System (fifth generation NVLink) |
| Grace CPU Architecture | Arm Neoverse V2, 72 cores per CPU |
| Grace CPU Memory | LPDDR5X with up to 480 GB/s per CPU |
| Cooling | Liquid cooling (direct liquid cooling required) |
| Form Factor | Rack-scale system (full rack) |
| Supported Precision Formats | FP4, FP8, FP16, BF16, TF32, FP32, INT8 |
| Confidential Computing | Supported (NVIDIA Hopper Confidential Computing architecture extended in Blackwell) |
| Software Stack | CUDA, cuDNN, TensorRT, NEMO, Triton Inference Server, NVIDIA AI Enterprise |
Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.
| Brand | NVIDIA |
| Category | GPUs |
| SKU | GB200-NVL36 |
| Part Number | GB200-NVL36 |
| Condition | New |
| Manufacturer Part Number | GB200-NVL36 |
| GPU Architecture | NVIDIA Blackwell |
| System Configuration | 36 × B200 GPUs + 18 × Grace CPUs (18 × GB200 Superchips) |
| Total HBM3e Memory | 6,912 GB |
| HBM3e Memory per GPU | 192 GB |
| HBM3e Memory Bandwidth per GPU | 8 TB/s |
| Total Memory Bandwidth (Rack) | Up to 288 TB/s aggregate |
| Peak AI Compute (FP4, Rack) | Up to 1.4 Exaflops |
| Peak AI Compute per GPU (FP4) | Up to 40 PFLOPS |
| NVLink All-to-All GPU Bandwidth | 1.8 TB/s (bidirectional) across all 36 GPUs |
| NVLink-C2C Bandwidth (Grace–Blackwell) | 900 GB/s bidirectional per Superchip |
| Interconnect Fabric | NVLink Switch System (fifth generation NVLink) |
| Grace CPU Architecture | Arm Neoverse V2, 72 cores per CPU |
| Grace CPU Memory | LPDDR5X with up to 480 GB/s per CPU |
| Cooling | Liquid cooling (direct liquid cooling required) |
| Form Factor | Rack-scale system (full rack) |
| Supported Precision Formats | FP4, FP8, FP16, BF16, TF32, FP32, INT8 |
| Confidential Computing | Supported (NVIDIA Hopper Confidential Computing architecture extended in Blackwell) |
| Software Stack | CUDA, cuDNN, TensorRT, NEMO, Triton Inference Server, NVIDIA AI Enterprise |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.