NVIDIA H100 NVL 188GB Dual-GPU

NVIDIA H100 NVL 188GB Dual-GPU

Brand: NVIDIA | Category: GPUs

SKU: 900-21010-0120-000 | Part #: 900-21010-0120-000 | MPN: 900-21010-0120-000

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the NVIDIA H100 NVL 188GB Dual-GPU

The NVIDIA H100 NVL 188GB Dual-GPU is a specialized accelerator pairing designed for large-scale language model inference and generative AI workloads. Each H100 NVL GPU features 141 streaming multiprocessors with 18,176 CUDA cores, 568 fourth-generation Tensor cores per SM, and 94GB of HBM3 memory per chip, delivering up to 60 teraflops of FP8 tensor performance per GPU. The dual-GPU configuration totals 188GB of memory and is optimized for inference at massive scale, leveraging Hopper architecture enhancements including Transformer Engine acceleration and Dynamic Tensor Memory optimization.

Designed for density-constrained environments where power efficiency and memory bandwidth are critical, the H100 NVL configuration enables sub-100W per-GPU thermal profiles while maintaining peak throughput for batched inference scenarios. The architecture supports 141GB/s aggregate memory bandwidth per GPU, enabling rapid token generation and multi-user inference serving across large parameter models. This configuration became available in Q2 2023 and has become foundational infrastructure for enterprise LLM deployment stacks.

The dual-GPU pairing is engineered for inference-optimized use cases where throughput per watt and memory-to-compute ratios directly impact deployment economics and model serving capacity.

Ideal for

  • Large language model inference serving with high concurrency and sub-100ms latency requirements
  • Multi-tenant generative AI API endpoints requiring per-request memory isolation and deterministic performance
  • Real-time retrieval-augmented generation (RAG) systems combining embedding and generation workloads
  • Batch inference pipelines for document analysis, code generation, and content summarization at enterprise scale
  • Fine-tuned model deployment where parameter efficiency meets throughput demands in resource-constrained data centers
  • Conversational AI and chatbot backend infrastructure supporting thousands of concurrent user sessions

Technical specifications

ManufacturerNVIDIA
ModelH100 NVL
ManufacturerPartNumber900-21010-0120-000
GPUCount2
TotalMemory188GB
MemoryPerGPU94GB HBM3
MemoryBandwidth282GB/s (141GB/s per GPU aggregate)
CUDACores36352
TensorCoresPerSM568
StreamingMultiprocessors282
ArchitectureNVIDIA Hopper
MaxPower250W
NVLinkBandwidth900GB/s (per connection)
FP8TensorPerformance120 teraflops (aggregate dual-GPU)
TF32TensorPerformance60 teraflops (aggregate dual-GPU)
TransformerEngineSupporttrue
LaunchQuarterQ2 2023
InterconnectSupportNVLink 4.0, PCIe Gen5

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandNVIDIA
CategoryGPUs
SKU900-21010-0120-000
Part Number900-21010-0120-000
ConditionNew
ModelH100 NVL
ManufacturerPartNumber900-21010-0120-000
GPUCount2
TotalMemory188GB
MemoryPerGPU94GB HBM3
MemoryBandwidth282GB/s (141GB/s per GPU aggregate)
CUDACores36352
TensorCoresPerSM568
StreamingMultiprocessors282
ArchitectureNVIDIA Hopper
MaxPower250W
NVLinkBandwidth900GB/s (per connection)
FP8TensorPerformance120 teraflops (aggregate dual-GPU)
TF32TensorPerformance60 teraflops (aggregate dual-GPU)
TransformerEngineSupporttrue
LaunchQuarterQ2 2023
InterconnectSupportNVLink 4.0, PCIe Gen5