PNY NVIDIA H100 NVL 94GB HBM3 PCIe Accelerator

PNY NVIDIA H100 NVL 94GB HBM3 PCIe Accelerator

Brand: NVIDIA | Category: GPUs

SKU: NVH100NVLTCGPU-KIT | Part #: NVH100NVLTCGPU-KIT | MPN: NVH100NVLTCGPU-KIT

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the PNY NVIDIA H100 NVL 94GB HBM3 PCIe Accelerator

The PNY NVIDIA H100 NVL 94GB HBM3 PCIe Accelerator (MPN: NVH100NVLTCGPU-KIT) is built on NVIDIA's Hopper GPU architecture and is specifically engineered for large language model (LLM) inference and training at scale. The H100 NVL variant ships in a dual-GPU NVLink configuration delivering a combined 188GB of HBM3 memory and 7.8 TB/s aggregate memory bandwidth, enabling the acceleration of models that cannot fit within the memory envelope of a single GPU. The PCIe Gen5 host interface makes this solution broadly compatible with modern enterprise server platforms without requiring proprietary interconnect fabrics at the host level.

At the silicon level, the H100 NVL leverages the GH100 die with fourth-generation Tensor Cores supporting FP8, FP16, BF16, TF32, and FP64 precision, alongside the Transformer Engine—an NVIDIA capability that dynamically selects precision per layer to maximize throughput without sacrificing model accuracy. Compared to prior-generation A100 PCIe solutions, the H100 NVL delivers substantially higher throughput for transformer-based inference workloads, making it a practical choice for production AI serving infrastructure. NVIDIA's Confidential Computing capability, also present on Hopper, allows tenant isolation and data-in-use protection in multi-tenant cloud and enterprise deployments.

As a PCIe form-factor card distributed by PNY, the NVH100NVLTCGPU-KIT targets enterprise data centers seeking Hopper-generation performance within standard OCP and proprietary rack environments. Launched in Q1 2024, the H100 NVL has experienced persistent supply constraints driven by acute global demand for LLM training and inference infrastructure. Organizations deploying this accelerator commonly pair it with NVIDIA's software stack including CUDA 12.x, cuDNN, TensorRT-LLM, and NEMO frameworks to realize full hardware capability across generative AI, scientific computing, and high-performance data analytics pipelines.

Ideal for

  • Large language model inference serving for production generative AI applications requiring high memory capacity and throughput across transformer models with tens to hundreds of billions of parameters
  • LLM fine-tuning and continued pre-training of foundation models where 94GB per-GPU HBM3 capacity (188GB across the NVLink pair) enables larger batch sizes and longer context windows
  • Enterprise retrieval-augmented generation (RAG) pipelines combining dense vector search with real-time LLM inference on a unified accelerator platform
  • High-performance scientific computing and simulation workloads in life sciences, energy, and financial modeling that benefit from FP64 double-precision throughput on Hopper Tensor Cores
  • Multi-tenant AI inference platforms in private cloud environments leveraging NVIDIA Confidential Computing on Hopper for workload isolation and data-in-use protection
  • Computer vision and multimodal model training combining large vision encoders with language decoders, where aggregate NVLink memory capacity reduces the need for model parallelism across additional nodes

Technical specifications

ManufacturerNVIDIA
BrandPNY (NVIDIA H100 NVL)
Manufacturer Part NumberNVH100NVLTCGPU-KIT
GPU ArchitectureNVIDIA Hopper (GH100)
GPU Memory per Card94 GB HBM3
GPU Memory ConfigurationDual-GPU NVLink pair — 188 GB HBM3 combined
Memory Bandwidth per Card3.9 TB/s
Aggregate NVLink Pair Memory Bandwidth7.8 TB/s
NVLink Interconnect Bandwidth600 GB/s bidirectional (NVLink 4.0)
Host InterfacePCIe Gen5 x16
Tensor Core Generation4th Generation (FP8, FP16, BF16, TF32, FP64)
Transformer EngineYes (2nd Generation)
Supported PrecisionsFP8, FP16, BF16, TF32, FP32, FP64, INT8
Confidential ComputingYes (NVIDIA Hopper Confidential Computing)
Form FactorPCIe Dual-slot (per GPU card in NVLink pair)
Thermal DesignPassive cooling (requires adequate chassis airflow)
Max Thermal Design Power (per GPU)400 W
ECC Memory SupportYes
CUDA Compute Capability9.0
Launch QuarterQ1 2024

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandNVIDIA
CategoryGPUs
SKUNVH100NVLTCGPU-KIT
Part NumberNVH100NVLTCGPU-KIT
ConditionNew
Manufacturer Part NumberNVH100NVLTCGPU-KIT
GPU ArchitectureNVIDIA Hopper (GH100)
GPU Memory per Card94 GB HBM3
GPU Memory ConfigurationDual-GPU NVLink pair — 188 GB HBM3 combined
Memory Bandwidth per Card3.9 TB/s
Aggregate NVLink Pair Memory Bandwidth7.8 TB/s
NVLink Interconnect Bandwidth600 GB/s bidirectional (NVLink 4.0)
Host InterfacePCIe Gen5 x16
Tensor Core Generation4th Generation (FP8, FP16, BF16, TF32, FP64)
Transformer EngineYes (2nd Generation)
Supported PrecisionsFP8, FP16, BF16, TF32, FP32, FP64, INT8
Confidential ComputingYes (NVIDIA Hopper Confidential Computing)
Form FactorPCIe Dual-slot (per GPU card in NVLink pair)
Thermal DesignPassive cooling (requires adequate chassis airflow)
Max Thermal Design Power (per GPU)400 W
ECC Memory SupportYes
CUDA Compute Capability9.0
Launch QuarterQ1 2024