Omnixon Global

NVIDIA H100 NVL 94GB HBM2e Dual-GPU PCIe Module

Brand: NVIDIA | Category: GPUs

SKU: 699-21010-0202-000 | Part #: 699-21010-0202-000 | MPN: 699-21010-0202-000

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the NVIDIA H100 NVL 94GB HBM2e Dual-GPU PCIe Module

The NVIDIA H100 NVL is a dual-GPU PCIe module featuring two H100 GPUs with 94GB of HBM2e memory per GPU (188GB total), designed for high-performance compute in standard enterprise server chassis without requiring SXM5 sockets. Each GPU delivers 141.1 TFLOPS of FP8 throughput and 70.6 TFLOPS of FP32, connected via NVLink for peer-to-peer communication and optimized for distributed inference workloads.

Purpose-built for large language model inferencing at scale, the H100 NVL addresses supply constraints affecting SXM variants while maintaining architectural feature parity including Transformer Engine support, fourth-generation NVLink, and PCIe 5.0 connectivity. The dual-GPU module fits into standard PCIe slots, enabling retrofit deployment in existing data center infrastructure without server redesign.

Target deployments span telecommunications, financial services, and public sector organizations requiring high-throughput text generation, reasoning tasks, and multi-concurrent model serving. The 188GB aggregate memory footprint supports inference of large foundation models with extended context windows and batch processing optimization.

Ideal for

  • Large language model inference serving (7B–70B+ parameter models) at production throughput in telecom and banking environments
  • Multi-model ensemble inference with concurrent request handling across customer-facing AI applications
  • Long-context document processing and retrieval-augmented generation (RAG) workloads requiring extended model context
  • Batch processing for financial risk modeling, fraud detection, and regulatory compliance analysis
  • Real-time conversational AI backends for customer service platforms and enterprise chat applications
  • Distributed inference pipelines leveraging NVLink inter-GPU communication for tensor parallelism

Technical specifications

ManufacturerNVIDIA
ManufacturerPartNumber699-21010-0202-000
ModelNameH100 NVL
FormFactorDual-GPU PCIe Module
GPUsPerModule2
MemoryPerGPU94GB HBM2e
TotalMemory188GB
MemoryBandwidth3.35TB/s per GPU
PCIeInterfacePCIe 5.0
NVLinkConnectivityFourth-generation NVLink (900GB/s per GPU)
FP32Performance70.6 TFLOPS per GPU
TensorFloat32Performance141.1 TFLOPS per GPU
FP8Performance565 TFLOPS per GPU
TransformerEngineSupported
CudaComputeCapability9.0
MaxPower700W per GPU
LaunchQuarter2023 Q4
TargetApplicationLLM Inference
ServerCompatibilityStandard PCIe x16 capable servers

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandNVIDIA
CategoryGPUs
SKU699-21010-0202-000
Part Number699-21010-0202-000
ConditionNew
ManufacturerPartNumber699-21010-0202-000
ModelNameH100 NVL
FormFactorDual-GPU PCIe Module
GPUsPerModule2
MemoryPerGPU94GB HBM2e
TotalMemory188GB
MemoryBandwidth3.35TB/s per GPU
PCIeInterfacePCIe 5.0
NVLinkConnectivityFourth-generation NVLink (900GB/s per GPU)
FP32Performance70.6 TFLOPS per GPU
TensorFloat32Performance141.1 TFLOPS per GPU
FP8Performance565 TFLOPS per GPU
TransformerEngineSupported
CudaComputeCapability9.0
MaxPower700W per GPU
LaunchQuarter2023 Q4
TargetApplicationLLM Inference
ServerCompatibilityStandard PCIe x16 capable servers