Brand: HPE | Category: GPUs
SKU: P40431-B21 | Part #: P40431-B21 | MPN: P40431-B21
Contact for Pricing — Request a Quote
The HPE NVIDIA A100 PCIe 40GB GPU Accelerator (part number P40431-B21) delivers 19.5 TFLOPS at FP32 precision, 312 TFLOPS at FP16, and 624 TOPS at INT8—the three critical compute tiers powering modern AI inference and training workloads. Built on NVIDIA's Ampere architecture, this accelerator combines 6,912 CUDA cores with 432 third-generation Tensor cores, backed by 40 GB of HBM2e memory running at 1,555 GB/s bandwidth. The card supports up to 7 Multi-Instance GPU (MIG) partitions, enabling cost-effective sharing across multiple workloads. With third-generation NVLink support (NVLink bridge required for multi-GPU configurations), ECC memory protection, and a 250 W thermal design power envelope, this HPE solution integrates seamlessly into HPE ProLiant DL380 Gen10 Plus and compatible HPE server platforms via PCIe Gen 4 x16 host interface in a full-height, full-length dual-slot form factor.
Enterprise AI infrastructure teams and machine learning operations groups deploy the A100 PCIe accelerator for large-scale deep learning, high-performance computing simulations, and real-time inference at scale. The 40 GB memory capacity eliminates costly data staging for many production AI pipelines, while Tensor Core performance—including 156 TFLOPS at TF32 and up to 624 TFLOPS at FP16 with sparsity enabled—addresses both dense and sparse model inference. Organizations standardizing on HPE infrastructure benefit from validated compatibility and unified support pathways across compute, storage, and networking.
Omnixon Global supplies the HPE NVIDIA A100 PCIe 40GB GPU Accelerator (P40431-B21) to enterprise customers across the UAE and Gulf region. To request a formal quotation, technical specifications, lead times, or volume pricing for this accelerator, contact Omnixon Global directly via our RFQ portal.
| Brand | HPE |
| Category | GPUs |
| SKU | P40431-B21 |
| Part Number | P40431-B21 |
| Condition | New |
| Manufacturer Part Number | P40431-B21 |
| GPU Architecture | NVIDIA Ampere |
| GPU Memory | 40 GB HBM2e |
| Memory Bandwidth | 1,555 GB/s |
| CUDA Cores | 6,912 |
| Tensor Cores | 432 (3rd Generation) |
| FP64 Peak Performance | 9.7 TFLOPS |
| FP32 Peak Performance | 19.5 TFLOPS |
| TF32 Tensor Core Peak Performance | 156 TFLOPS (312 TFLOPS with sparsity) |
| FP16 Tensor Core Peak Performance | 312 TFLOPS (624 TFLOPS with sparsity) |
| INT8 Tensor Core Peak Performance | 624 TOPS (1,248 TOPS with sparsity) |
| Host Interface | PCIe Gen 4 x16 |
| Form Factor | Full-height, full-length (FHFL) dual-slot |
| Thermal Design Power (TDP) | 250 W |
| Multi-Instance GPU (MIG) | Up to 7 MIG instances |
| NVLink | Third-generation NVLink (NVLink bridge required for multi-GPU) |
| ECC Memory | Yes |
| Compatible HPE Platforms | HPE ProLiant DL380 Gen10 Plus and compatible HPE server platforms |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.