Brand: HPE | Category: GPUs
SKU: P40429-B21 | Part #: P40429-B21 | MPN: P40429-B21
Contact for Pricing — Request a Quote
The NVIDIA A100 40GB PCIe Gen4 GPU for HPE ProLiant (P40429-B21) is built on NVIDIA's Ampere architecture, delivering third-generation Tensor Core technology that accelerates mixed-precision AI training and inference workloads at datacenter scale. The 40GB HBM2e memory subsystem provides 1,555 GB/s of memory bandwidth, enabling large-model training and high-throughput inferencing that would otherwise require multiple previous-generation accelerators. PCIe Gen4 host connectivity ensures the card integrates cleanly into HPE ProLiant server platforms without the NVLink or NVSwitch fabric requirements of SXM form-factor variants, simplifying rack-level deployment.
Designed as a passive-cooled accelerator, the A100 PCIe relies on the managed airflow infrastructure of HPE ProLiant rack servers, making it well-suited for controlled datacenter environments where chassis-level cooling is already optimized for high-density compute. The card's 300W TDP is handled entirely through system airflow, eliminating the noise and mechanical complexity of onboard fans while maintaining full sustained performance under continuous workload. Multi-Instance GPU (MIG) technology allows the A100 to be partitioned into as many as seven isolated GPU instances, each with dedicated memory, cache, and compute resources, giving operators fine-grained control over workload isolation and resource utilization in shared or multi-tenant environments.
As an HPE factory-integrated option, P40429-B21 is validated against HPE ProLiant Gen10 Plus and Gen11 server platforms, with HPE iLO-based manageability and compatibility with HPE's software ecosystem including support through HPE Pointnext services infrastructure. The combination of Ampere Tensor Cores, structural sparsity acceleration, and NVLink-capable die (in multi-GPU PCIe topologies via NVLink Bridge where supported) positions this GPU for demanding workloads spanning HPC simulation, large language model development, and real-time analytics pipelines.
| Manufacturer | HPE |
| Manufacturer Part Number | P40429-B21 |
| GPU Architecture | NVIDIA Ampere |
| GPU Memory | 40 GB HBM2e |
| Memory Bandwidth | 1,555 GB/s |
| Host Interface | PCIe Gen4 x16 |
| Form Factor | Full-height, full-length (FHFL) dual-slot |
| Cooling | Passive (system airflow dependent) |
| Thermal Design Power (TDP) | 300 W |
| FP64 (Double Precision) Peak Performance | 9.7 TFLOPS |
| FP32 (Single Precision) Peak Performance | 19.5 TFLOPS |
| TF32 Tensor Core Peak Performance | 156 TFLOPS (312 TFLOPS with sparsity) |
| FP16 Tensor Core Peak Performance | 312 TFLOPS (624 TFLOPS with sparsity) |
| INT8 Tensor Core Peak Performance | 624 TOPS (1,248 TOPS with sparsity) |
| Multi-Instance GPU (MIG) | Up to 7 MIG instances (1g.5gb profile) |
| NVLink Support | 600 GB/s bidirectional (via NVLink Bridge, two-GPU PCIe configurations) |
| ECC Memory | Yes |
| Supported Server Platforms | HPE ProLiant Gen10 Plus and Gen11 series |
| GPU Direct RDMA | Supported |
| NVIDIA CUDA Compute Capability | 8.0 |
Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.
| Brand | HPE |
| Category | GPUs |
| SKU | P40429-B21 |
| Part Number | P40429-B21 |
| Condition | New |
| Manufacturer Part Number | P40429-B21 |
| GPU Architecture | NVIDIA Ampere |
| GPU Memory | 40 GB HBM2e |
| Memory Bandwidth | 1,555 GB/s |
| Host Interface | PCIe Gen4 x16 |
| Form Factor | Full-height, full-length (FHFL) dual-slot |
| Cooling | Passive (system airflow dependent) |
| Thermal Design Power (TDP) | 300 W |
| FP64 (Double Precision) Peak Performance | 9.7 TFLOPS |
| FP32 (Single Precision) Peak Performance | 19.5 TFLOPS |
| TF32 Tensor Core Peak Performance | 156 TFLOPS (312 TFLOPS with sparsity) |
| FP16 Tensor Core Peak Performance | 312 TFLOPS (624 TFLOPS with sparsity) |
| INT8 Tensor Core Peak Performance | 624 TOPS (1,248 TOPS with sparsity) |
| Multi-Instance GPU (MIG) | Up to 7 MIG instances (1g.5gb profile) |
| NVLink Support | 600 GB/s bidirectional (via NVLink Bridge, two-GPU PCIe configurations) |
| ECC Memory | Yes |
| Supported Server Platforms | HPE ProLiant Gen10 Plus and Gen11 series |
| GPU Direct RDMA | Supported |
| NVIDIA CUDA Compute Capability | 8.0 |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.