NVIDIA H100 NVL 94GB HBM2e Dual-GPU

NVIDIA H100 NVL 94GB HBM2e Dual-GPU

Brand: NVIDIA | Category: GPUs

SKU: 900-21010-0030-000 | Part #: 900-21010-0030-000 | MPN: 900-21010-0030-000

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the NVIDIA H100 NVL 94GB HBM2e Dual-GPU

The NVIDIA H100 NVL 94GB HBM2e Dual-GPU is a high-capacity accelerator module built on the NVIDIA Hopper architecture, engineered specifically for large-scale generative AI inference, training, and high-performance computing workloads in enterprise datacenter environments. Featuring dual H100 GPUs on a single NVLink board, the platform delivers a combined 94GB of HBM2e memory, enabling organizations to run exceptionally large language models and multi-modal AI workloads entirely within GPU memory without model partitioning overhead.

The H100 NVL leverages the fourth-generation NVLink interconnect to bind the two GPUs into a unified, high-bandwidth memory domain, providing up to 900 GB/s of GPU-to-GPU bandwidth. The Hopper architecture introduces the Transformer Engine, which dynamically applies FP8 and FP16 mixed-precision computation to dramatically accelerate transformer-based model inference while sustaining accuracy. Fourth-generation Tensor Cores further extend throughput across FP8, FP16, BF16, TF32, and INT8 precisions, making the platform well-suited across the full AI development lifecycle from exploratory research to production serving.

The 900-21010-0030-000 SKU is designed for PCIe-based server integration, offering broad compatibility with standard enterprise rack infrastructure without requiring proprietary NVLink Switch fabric chassis. This positions the H100 NVL as a practical choice for organizations scaling AI infrastructure incrementally, with each dual-GPU board delivering the combined compute density of two discrete H100 SXM5-class processors in a single PCIe form factor, supported by NVIDIA's mature software ecosystem including CUDA, cuDNN, TensorRT, and the full NGC catalog of enterprise-ready AI frameworks and containers.

Ideal for

  • Large language model inference serving for enterprise generative AI applications requiring full model residency in GPU memory, including 70B+ parameter transformer models
  • Distributed deep learning training for vision, language, and multimodal foundation models where high inter-GPU bandwidth reduces gradient synchronization overhead
  • High-throughput scientific simulation and HPC workloads in computational chemistry, climate modeling, and genomics that benefit from large unified GPU memory capacity
  • Enterprise AI platform consolidation, enabling multiple independent AI workloads to run concurrently via MIG (Multi-Instance GPU) partitioning across the dual-GPU module
  • Retrieval-augmented generation (RAG) pipelines and vector database acceleration where simultaneous embedding, retrieval, and generation tasks demand sustained memory bandwidth
  • Real-time AI inference at scale for financial risk modeling, fraud detection, and recommendation engines requiring low-latency, high-throughput GPU compute in standard PCIe datacenter servers

Technical specifications

ManufacturerNVIDIA
Manufacturer Part Number900-21010-0030-000
Product LineNVIDIA H100 NVL
ArchitectureNVIDIA Hopper (GH100)
GPU ConfigurationDual-GPU (2× H100 per board)
Total GPU Memory94 GB HBM2e
Memory per GPU47 GB HBM2e
Memory Bandwidth (Total)Up to 7.8 TB/s aggregate
NVLink Interconnect4th Generation NVLink, up to 900 GB/s GPU-to-GPU bandwidth
FP8 Tensor Core PerformanceUp to 3,958 TFLOPS (sparse) per board
FP16 Tensor Core PerformanceUp to 1,979 TFLOPS (sparse) per board
TF32 Tensor Core PerformanceUp to 990 TFLOPS (sparse) per board
Form FactorPCIe (dual-slot)
Host InterfacePCIe Gen 5 x16
Transformer EngineYes (FP8 and FP16 mixed precision)
Multi-Instance GPU (MIG)Supported
Confidential ComputingSupported
ECC Memory SupportYes
Total Board Power (TDP)350 W
Operating Temperature0°C to 35°C (inlet)

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandNVIDIA
CategoryGPUs
SKU900-21010-0030-000
Part Number900-21010-0030-000
ConditionNew
Manufacturer Part Number900-21010-0030-000
Product LineNVIDIA H100 NVL
ArchitectureNVIDIA Hopper (GH100)
GPU ConfigurationDual-GPU (2× H100 per board)
Total GPU Memory94 GB HBM2e
Memory per GPU47 GB HBM2e
Memory Bandwidth (Total)Up to 7.8 TB/s aggregate
NVLink Interconnect4th Generation NVLink, up to 900 GB/s GPU-to-GPU bandwidth
FP8 Tensor Core PerformanceUp to 3,958 TFLOPS (sparse) per board
FP16 Tensor Core PerformanceUp to 1,979 TFLOPS (sparse) per board
TF32 Tensor Core PerformanceUp to 990 TFLOPS (sparse) per board
Form FactorPCIe (dual-slot)
Host InterfacePCIe Gen 5 x16
Transformer EngineYes (FP8 and FP16 mixed precision)
Multi-Instance GPU (MIG)Supported
Confidential ComputingSupported
ECC Memory SupportYes
Total Board Power (TDP)350 W
Operating Temperature0°C to 35°C (inlet)

Frequently Asked Questions about NVIDIA H100 NVL 94GB HBM2e Dual-GPU

What server platforms accept the NVIDIA H100 NVL 94GB HBM2e Dual-GPU?

Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.

How long is the lead time on AI GPUs?

Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.

Do you supply matched networking (Quantum InfiniBand / Spectrum-X)?

Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.

Can you help with NVIDIA AI Enterprise licensing?

Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.