Brand: NVIDIA | Category: GPUs
SKU: GB200-NVL2 | Part #: GB200-NVL2 | MPN: GB200-NVL2
Contact for Pricing — Request a Quote
The NVIDIA GB200 NVL2 Grace Blackwell Superchip is a tightly integrated multi-chip module combining two NVIDIA Blackwell B200 GPU dies with a single NVIDIA Grace CPU via an ultra-high-bandwidth NVLink-C2C interconnect. The Grace CPU is an Arm Neoverse V2-based processor delivering 72 high-performance cores, while the two Blackwell GPU dies together present a unified 208 billion transistor compute engine. The NVLink-C2C chip-to-chip interconnect provides 900 GB/s of total bidirectional bandwidth between CPU and GPU, eliminating PCIe bottlenecks and enabling coherent memory access across a combined CPU and GPU HBM3e memory pool.
The GB200 NVL2 is engineered specifically for the most demanding generative AI training and inference workloads at scale. Each Blackwell GPU die incorporates a fifth-generation NVLink fabric interface, second-generation Transformer Engine with support for FP4 and FP8 precision, and a dedicated in-fabric all-reduce engine for accelerated collective communications. The aggregate HBM3e memory capacity across both GPU dies reaches 192 GB with 8 TB/s of total memory bandwidth, enabling large language models and multimodal foundation models to reside entirely in GPU-accessible high-bandwidth memory without host offloading.
Designed for dense rack-scale AI infrastructure, the GB200 NVL2 operates within the NVIDIA MGX mechanical specification and is the core building block of NVIDIA GB200 NVL72 rack-scale systems. It supports liquid cooling to manage the high thermal envelope associated with sustained AI compute throughput. Enterprise data centers deploying GB200 NVL2 gain access to NVIDIA's full software ecosystem including CUDA, cuDNN, TensorRT-LLM, and NeMo frameworks, making it suitable for sovereign AI deployments, large-scale model training clusters, and high-throughput inference services across UAE, GCC, EMEA, and APAC regions.
| Manufacturer | NVIDIA |
| Manufacturer Part Number | GB200-NVL2 |
| Product Family | Grace Blackwell Superchip |
| GPU Architecture | NVIDIA Blackwell |
| CPU Architecture | NVIDIA Grace (Arm Neoverse V2, 72 cores) |
| GPU Dies per Module | 2x Blackwell B200 GPU dies |
| Total Transistors | 208 billion (across both Blackwell GPU dies) |
| GPU Memory | 192 GB HBM3e (aggregate across both GPU dies) |
| GPU Memory Bandwidth | 8 TB/s (aggregate across both GPU dies) |
| CPU-GPU Interconnect | NVLink-C2C, 900 GB/s total bidirectional bandwidth |
| NVLink Generation | 5th Generation NVLink |
| Tensor Core Generation | 5th Generation (supports FP4, FP8, FP16, BF16, TF32, FP64) |
| Transformer Engine | 2nd Generation |
| System Memory (CPU) | Up to 480 GB LPDDR5X (Grace CPU) |
| System Memory Bandwidth (CPU) | Up to 546 GB/s (Grace LPDDR5X) |
| Form Factor | NVIDIA MGX Superchip Module (NVL2 configuration) |
| Cooling | Liquid cooling required |
| Interconnect Fabric Role | Base building block for GB200 NVL72 rack-scale systems |
| Software Ecosystem | CUDA, cuDNN, TensorRT-LLM, NeMo, NVIDIA AI Enterprise |
Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.
| Brand | NVIDIA |
| Category | GPUs |
| SKU | GB200-NVL2 |
| Part Number | GB200-NVL2 |
| Condition | New |
| Manufacturer Part Number | GB200-NVL2 |
| Product Family | Grace Blackwell Superchip |
| GPU Architecture | NVIDIA Blackwell |
| CPU Architecture | NVIDIA Grace (Arm Neoverse V2, 72 cores) |
| GPU Dies per Module | 2x Blackwell B200 GPU dies |
| Total Transistors | 208 billion (across both Blackwell GPU dies) |
| GPU Memory | 192 GB HBM3e (aggregate across both GPU dies) |
| GPU Memory Bandwidth | 8 TB/s (aggregate across both GPU dies) |
| CPU-GPU Interconnect | NVLink-C2C, 900 GB/s total bidirectional bandwidth |
| NVLink Generation | 5th Generation NVLink |
| Tensor Core Generation | 5th Generation (supports FP4, FP8, FP16, BF16, TF32, FP64) |
| Transformer Engine | 2nd Generation |
| System Memory (CPU) | Up to 480 GB LPDDR5X (Grace CPU) |
| System Memory Bandwidth (CPU) | Up to 546 GB/s (Grace LPDDR5X) |
| Form Factor | NVIDIA MGX Superchip Module (NVL2 configuration) |
| Cooling | Liquid cooling required |
| Interconnect Fabric Role | Base building block for GB200 NVL72 rack-scale systems |
| Software Ecosystem | CUDA, cuDNN, TensorRT-LLM, NeMo, NVIDIA AI Enterprise |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.