NVIDIA GB200 NVL72 Compute Blade

NVIDIA GB200 NVL72 Compute Blade

Brand: Dell | Category: GPUs

SKU: 900-5G200-0000-000 | Part #: 900-5G200-0000-000 | MPN: 900-5G200-0000-000

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the NVIDIA GB200 NVL72 Compute Blade

The NVIDIA GB200 NVL72 Compute Blade, manufactured by Dell, is a high-density liquid-cooled compute platform built around NVIDIA's Blackwell architecture. The NVL72 configuration integrates 36 Grace CPU modules and 72 Blackwell B200 GPU dies into a single NVLink-connected rack-scale unit, enabling all 72 GPUs to operate as a unified, high-bandwidth compute fabric. The Grace CPU component is based on the Arm Neoverse V2 architecture, while each B200 GPU delivers second-generation Transformer Engine support and fifth-generation NVLink interconnects running at 1.8 TB/s of bisectional bandwidth across the full 72-GPU domain.

Designed for the most demanding AI infrastructure deployments, the GB200 NVL72 delivers up to 1.4 exaflops of AI compute (FP4 precision) across the full rack unit, with a substantial uplift in FP8 and FP16 throughput for training large-scale foundation models. The integrated NVLink Switch fabric eliminates PCIe bottlenecks between GPUs, enabling low-latency all-reduce operations critical for distributed training. The platform uses a direct liquid cooling (DLC) thermal solution, which is mandatory given the rack-level power envelope, and is designed for deployment in data centers equipped with rear-door heat exchangers or direct liquid cooling infrastructure.

The Dell GB200 NVL72 Compute Blade carries the manufacturer part number 900-5G200-0000-000 and is targeted at enterprise data centers, hyperscale operators, national AI research facilities, and sovereign cloud deployments across regions including the UAE, GCC, EMEA, and APAC. It is positioned as a rack-scale AI supercomputer building block, particularly suited for generative AI training, large language model inference at scale, scientific simulation, and high-performance computing workloads that demand extreme memory bandwidth and multi-GPU coherency.

Ideal for

  • Large language model (LLM) pre-training and fine-tuning at foundation model scale, leveraging the unified 72-GPU NVLink fabric for efficient distributed gradient synchronization
  • Generative AI inference serving for enterprise applications requiring high throughput and low latency across concurrent user requests at data center scale
  • High-performance scientific computing and simulation workloads in fields such as computational fluid dynamics, molecular dynamics, and climate modeling that require tightly coupled multi-GPU execution
  • Sovereign AI infrastructure buildouts for government, national research institutions, and regulated industries requiring on-premises, high-density AI compute within UAE, GCC, EMEA, and APAC jurisdictions
  • Multimodal AI model development including vision-language and video generation models that demand large aggregate GPU memory capacity across a coherent memory domain
  • Enterprise data center consolidation for AI and HPC, replacing multi-rack GPU clusters with a single rack-scale NVL72 unit to reduce floor space, networking complexity, and operational overhead

Technical specifications

ManufacturerDell
ModelNVIDIA GB200 NVL72 Compute Blade
Manufacturer Part Number900-5G200-0000-000
GPU ArchitectureNVIDIA Blackwell
GPU Configuration72x NVIDIA B200 GPU dies (36 GB200 Grace-Blackwell Superchips)
CPU Configuration36x NVIDIA Grace CPUs (Arm Neoverse V2, 72 cores each)
Total GPUs per NVL7272
Total Grace CPUs per NVL7236
GPU InterconnectFifth-generation NVLink with NVLink Switch, 1.8 TB/s bisectional bandwidth
AI Compute Performance (FP4)Up to 1.4 exaflops per NVL72 rack unit
GPU Memory TypeHBM3e
GPU Memory per B200 Die192 GB
Total HBM3e Memory (NVL72)13.5 TB
GPU Memory Bandwidth per B2008 TB/s
Thermal SolutionDirect Liquid Cooling (DLC)
Form FactorRack-scale NVLink rack unit (full rack)
NVLink Switch GenerationNVLink Switch Gen 5
Target WorkloadsGenerative AI training, LLM inference, HPC, scientific simulation
Supported Precision FormatsFP4, FP8, FP16, BF16, TF32, FP32, INT8
Transformer Engine GenerationSecond-generation NVIDIA Transformer Engine
Host InterfacePCIe Gen 5 (Grace CPU to B200 GPU via NVLink-C2C within each Superchip)

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandDell
CategoryGPUs
SKU900-5G200-0000-000
Part Number900-5G200-0000-000
ConditionNew
ModelNVIDIA GB200 NVL72 Compute Blade
Manufacturer Part Number900-5G200-0000-000
GPU ArchitectureNVIDIA Blackwell
GPU Configuration72x NVIDIA B200 GPU dies (36 GB200 Grace-Blackwell Superchips)
CPU Configuration36x NVIDIA Grace CPUs (Arm Neoverse V2, 72 cores each)
Total GPUs per NVL7272
Total Grace CPUs per NVL7236
GPU InterconnectFifth-generation NVLink with NVLink Switch, 1.8 TB/s bisectional bandwidth
AI Compute Performance (FP4)Up to 1.4 exaflops per NVL72 rack unit
GPU Memory TypeHBM3e
GPU Memory per B200 Die192 GB
Total HBM3e Memory (NVL72)13.5 TB
GPU Memory Bandwidth per B2008 TB/s
Thermal SolutionDirect Liquid Cooling (DLC)
Form FactorRack-scale NVLink rack unit (full rack)
NVLink Switch GenerationNVLink Switch Gen 5
Target WorkloadsGenerative AI training, LLM inference, HPC, scientific simulation
Supported Precision FormatsFP4, FP8, FP16, BF16, TF32, FP32, INT8
Transformer Engine GenerationSecond-generation NVIDIA Transformer Engine
Host InterfacePCIe Gen 5 (Grace CPU to B200 GPU via NVLink-C2C within each Superchip)

Frequently Asked Questions about NVIDIA GB200 NVL72 Compute Blade

What server platforms accept the NVIDIA GB200 NVL72 Compute Blade?

Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.

How long is the lead time on AI GPUs?

Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.

Do you supply matched networking (Quantum InfiniBand / Spectrum-X)?

Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.

Can you help with NVIDIA AI Enterprise licensing?

Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.