Brand: Gigabyte | Category: GPUs
SKU: 900-21010-0010-000 | Part #: 900-21010-0010-000 | MPN: 900-21010-0010-000
Contact for Pricing — Request a Quote
The NVIDIA H100 NVL 94GB HBM2e Dual-GPU Module, manufactured by Gigabyte under part number 900-21010-0010-000, is a high-density dual-GPU solution built on NVIDIA's Hopper architecture. The module integrates two H100 GPUs interconnected via NVLink, delivering a combined 94GB of HBM2e memory and exceptional aggregate memory bandwidth, making it one of the most capable accelerator platforms available for large-scale AI inference and training workloads. The Hopper architecture introduces the Transformer Engine with FP8 precision support, enabling dramatic throughput gains for large language models and generative AI applications compared to prior GPU generations.
Designed specifically for demanding datacenter environments, this dual-GPU module supports PCIe Gen 5 host connectivity and is engineered to fit into standard enterprise server chassis that accommodate the NVL form factor. The combined NVLink interconnect between the two H100 dies enables GPU-to-GPU communication at bandwidths far exceeding what PCIe alone can provide, which is critical for model-parallel and tensor-parallel inference workloads where data must move rapidly between GPU partitions. The module supports NVIDIA's full software stack, including CUDA, TensorRT, Triton Inference Server, and the NCCL communications library.
Targeted at enterprises operating AI factories, HPC clusters, and large-scale inferencing platforms across regions including the UAE, GCC, EMEA, and APAC, the H100 NVL 94GB module from Gigabyte provides a production-grade, datacenter-validated solution for organizations deploying frontier AI models, scientific simulation workloads, and high-throughput data analytics pipelines. Its high memory capacity and bandwidth make it particularly well suited to serving very large transformer-based models in a single module without model sharding across multiple nodes.
| Manufacturer | Gigabyte |
| Manufacturer Part Number | 900-21010-0010-000 |
| GPU Architecture | NVIDIA Hopper (H100) |
| Module Configuration | Dual-GPU (NVL form factor) |
| Total GPU Memory | 94 GB HBM2e |
| Memory Type | HBM2e |
| GPU Interconnect | NVLink (fourth generation) |
| Host Interface | PCIe Gen 5 |
| FP8 Tensor Core Performance | Supported via Transformer Engine |
| Supported Precisions | FP64, FP32, TF32, BF16, FP16, FP8, INT8 |
| CUDA Compute Capability | 9.0 |
| Form Factor | NVL Dual-GPU Module |
| Cooling | Passive (requires system-level airflow) |
| Software Stack | CUDA, TensorRT, Triton Inference Server, NCCL, cuDNN |
| Operating System Support | Linux (CUDA-compatible distributions) |
| Target Deployment | Enterprise datacenter, AI inference, HPC |
Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.
| Brand | Gigabyte |
| Category | GPUs |
| SKU | 900-21010-0010-000 |
| Part Number | 900-21010-0010-000 |
| Condition | New |
| Manufacturer Part Number | 900-21010-0010-000 |
| GPU Architecture | NVIDIA Hopper (H100) |
| Module Configuration | Dual-GPU (NVL form factor) |
| Total GPU Memory | 94 GB HBM2e |
| Memory Type | HBM2e |
| GPU Interconnect | NVLink (fourth generation) |
| Host Interface | PCIe Gen 5 |
| FP8 Tensor Core Performance | Supported via Transformer Engine |
| Supported Precisions | FP64, FP32, TF32, BF16, FP16, FP8, INT8 |
| CUDA Compute Capability | 9.0 |
| Form Factor | NVL Dual-GPU Module |
| Cooling | Passive (requires system-level airflow) |
| Software Stack | CUDA, TensorRT, Triton Inference Server, NCCL, cuDNN |
| Operating System Support | Linux (CUDA-compatible distributions) |
| Target Deployment | Enterprise datacenter, AI inference, HPC |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.