Lenovo ThinkSystem SR680a V3 with NVIDIA GB200 NVL72 Rack-Scale System

Lenovo ThinkSystem SR680a V3 with NVIDIA GB200 NVL72 Rack-Scale System

Brand: Lenovo | Category: GPUs

SKU: 7DHD (GB200 NVL72 rack config) | Part #: 7DHD (GB200 NVL72 rack config) | MPN: 7DHD (GB200 NVL72 rack config)

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the Lenovo ThinkSystem SR680a V3 with NVIDIA GB200 NVL72 Rack-Scale System

The Lenovo ThinkSystem SR680a V3 with NVIDIA GB200 NVL72 is a rack-scale AI supercomputing system engineered for the most demanding large-scale generative AI training, inference, and HPC workloads. Built around NVIDIA's GB200 NVL72 architecture, the system integrates 36 Grace CPU modules and 72 Blackwell B200 GPU dies within a single NVLink-interconnected rack domain, delivering unified high-bandwidth memory and compute fabric at unprecedented scale. The GB200 NVL72 configuration treats the entire rack as a single logical GPU, connected via fifth-generation NVLink with an aggregate NVLink bandwidth of 130 TB/s across the system, enabling model parallelism strategies that were previously impractical within a single rack enclosure.

The system is purpose-built for trillion-parameter foundation model training and real-time inference, leveraging NVIDIA Blackwell GPU architecture with support for FP4, FP8, FP16, BF16, TF32, and FP64 precision formats. Each B200 GPU die delivers up to 20 petaflops of AI compute at FP4 precision, and the 72-GPU NVL72 configuration aggregates this to rack-scale throughput suited for next-generation large language models, multimodal AI, scientific simulation, and drug discovery pipelines. Liquid cooling is integral to the design, with Lenovo's direct liquid cooling infrastructure supporting the extreme thermal demands of the Grace Blackwell Superchip modules throughout the rack.

Lenovo integrates the GB200 NVL72 into the ThinkSystem SR680a V3 platform with enterprise-grade systems management via Lenovo XClarity Administrator, alongside support for the NVIDIA AI Enterprise software stack, CUDA, cuDNN, and the NVIDIA NeMo framework. The system is designed for deployment in hyperscale and enterprise AI data centers requiring maximum GPU memory bandwidth, NVLink fabric scale-out, and dense compute per rack unit, making it a flagship platform for organizations building sovereign AI infrastructure across EMEA, GCC, UAE, and APAC regions.

Ideal for

  • Large-scale foundation model and trillion-parameter LLM training requiring rack-scale NVLink-connected GPU memory pools
  • High-throughput generative AI inference serving for real-time applications including conversational AI and multimodal workloads
  • Computational drug discovery and molecular dynamics simulation leveraging FP64 and mixed-precision Blackwell compute
  • Sovereign AI infrastructure build-out for national AI initiatives, government agencies, and hyperscale cloud operators in EMEA and GCC
  • Digital twin simulation and large-scale scientific HPC workloads requiring dense GPU interconnect and high-bandwidth memory
  • Enterprise AI platform consolidation enabling simultaneous multi-tenant LLM fine-tuning and inference on shared rack-scale GPU fabric

Technical specifications

ManufacturerLenovo
Product LineThinkSystem SR680a V3
Manufacturer Part Number7DHD (GB200 NVL72 rack configuration)
GPU ArchitectureNVIDIA Blackwell (GB200)
System ConfigurationGB200 NVL72 — 36 Grace CPUs + 72 Blackwell B200 GPU dies in rack-scale NVLink domain
GPU Count per Rack72 x NVIDIA B200 GPU dies (via 36 GB200 Grace Blackwell Superchips)
CPU36 x NVIDIA Grace CPU (72-core Arm Neoverse V2 per module)
GPU Memory8,640 GB HBM3e total (120 GB HBM3e per B200 GPU die)
GPU Memory BandwidthUp to 8.0 TB/s per B200 GPU die; aggregate rack-scale bandwidth across NVLink fabric
NVLink InterconnectNVLink 5 — 130 TB/s aggregate bidirectional NVLink bandwidth across full NVL72 rack
AI Compute (FP4)Up to 20 petaFLOPS per B200 GPU die (1,440 petaFLOPS aggregate across 72 GPUs)
Supported Precision FormatsFP4, FP8, FP16, BF16, TF32, FP64, INT8
CoolingDirect liquid cooling (DLC) — mandatory for GB200 NVL72 thermal envelope
Form FactorRack-scale system (full rack enclosure)
Network ConnectivityNVIDIA Quantum-X800 InfiniBand or Spectrum-X800 Ethernet (scale-out fabric)
Systems ManagementLenovo XClarity Administrator; NVIDIA Base Command Manager compatible
Software StackNVIDIA AI Enterprise, CUDA, cuDNN, NCCL, NVIDIA NeMo, Triton Inference Server
Operating System SupportUbuntu Linux, Red Hat Enterprise Linux (RHEL)
Target DeploymentHyperscale and enterprise AI data centers; sovereign AI infrastructure

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandLenovo
CategoryGPUs
SKU7DHD (GB200 NVL72 rack config)
Part Number7DHD (GB200 NVL72 rack config)
ConditionNew
Product LineThinkSystem SR680a V3
Manufacturer Part Number7DHD (GB200 NVL72 rack configuration)
GPU ArchitectureNVIDIA Blackwell (GB200)
System ConfigurationGB200 NVL72 — 36 Grace CPUs + 72 Blackwell B200 GPU dies in rack-scale NVLink domain
GPU Count per Rack72 x NVIDIA B200 GPU dies (via 36 GB200 Grace Blackwell Superchips)
CPU36 x NVIDIA Grace CPU (72-core Arm Neoverse V2 per module)
GPU Memory8,640 GB HBM3e total (120 GB HBM3e per B200 GPU die)
GPU Memory BandwidthUp to 8.0 TB/s per B200 GPU die; aggregate rack-scale bandwidth across NVLink fabric
NVLink InterconnectNVLink 5 — 130 TB/s aggregate bidirectional NVLink bandwidth across full NVL72 rack
AI Compute (FP4)Up to 20 petaFLOPS per B200 GPU die (1,440 petaFLOPS aggregate across 72 GPUs)
Supported Precision FormatsFP4, FP8, FP16, BF16, TF32, FP64, INT8
CoolingDirect liquid cooling (DLC) — mandatory for GB200 NVL72 thermal envelope
Form FactorRack-scale system (full rack enclosure)
Network ConnectivityNVIDIA Quantum-X800 InfiniBand or Spectrum-X800 Ethernet (scale-out fabric)
Systems ManagementLenovo XClarity Administrator; NVIDIA Base Command Manager compatible
Software StackNVIDIA AI Enterprise, CUDA, cuDNN, NCCL, NVIDIA NeMo, Triton Inference Server
Operating System SupportUbuntu Linux, Red Hat Enterprise Linux (RHEL)
Target DeploymentHyperscale and enterprise AI data centers; sovereign AI infrastructure

Frequently Asked Questions about Lenovo ThinkSystem SR680a V3 with NVIDIA GB200 NVL72 Rack-Scale System

What server platforms accept the Lenovo ThinkSystem SR680a V3 with NVIDIA GB200 NVL72 Rack-Scale System?

Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.

How long is the lead time on AI GPUs?

Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.

Do you supply matched networking (Quantum InfiniBand / Spectrum-X)?

Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.

Can you help with NVIDIA AI Enterprise licensing?

Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.