Brand: HPE | Category: GPUs
SKU: P40435-B21 | Part #: P40435-B21 | MPN: P40435-B21
Contact for Pricing — Request a Quote
The HPE NVIDIA A40 48GB PCIe GPU Accelerator delivers 346 TFLOPS at FP32, 692 TFLOPS at BF16, and 1,384 TFLOPS at FP8/INT8, positioning it as a versatile compute platform for mixed-precision AI and professional visualization workloads. Built on NVIDIA's Ampere architecture, this accelerator combines 10,752 CUDA cores with 336 third-generation Tensor Cores and 84 second-generation RT Cores, all backed by 48 GB of ECC-protected GDDR6 memory on a 384-bit bus capable of 696 GB/s throughput. The passive-cooled design operates at 300 W TDP and connects via PCIe 4.0 x16, making it suitable for deployment in HPE ProLiant DL380 Gen10 Plus, DL360 Gen10 Plus, and compatible Apollo systems.
For infrastructure teams evaluating GPU acceleration, the A40 offers enterprise-grade features including NVLink support for direct peer-to-peer communication between up to two GPUs, full ECC memory protection for mission-critical compute, and NVIDIA vGPU support for flexible workload sharing. The full-height, full-length dual-slot form factor integrates seamlessly into HPE server environments, while native support for CUDA, OpenCL, Vulkan, OpenGL, and DirectX 12 Ultimate ensures broad software compatibility. Four DisplayPort 1.4 outputs enable multi-monitor visualization workflows without additional adapters. Part number P40435-B21 identifies this HPE-manufactured module for procurement accuracy. AI infrastructure teams, data center architects, and enterprises requiring balanced performance in both training and inference scenarios will find the A40's memory capacity and bandwidth instrumental for production deployments.
To request specifications, pricing, or inventory availability for the HPE NVIDIA A40 (P40435-B21), contact Omnixon Global's enterprise solutions team today.
| Brand | HPE |
| Category | GPUs |
| SKU | P40435-B21 |
| Part Number | P40435-B21 |
| Condition | New |
| Manufacturer Part Number | P40435-B21 |
| GPU Architecture | NVIDIA Ampere |
| GPU Memory | 48 GB GDDR6 ECC |
| Memory Bus Width | 384-bit |
| Memory Bandwidth | 696 GB/s |
| CUDA Cores | 10752 |
| Tensor Cores | 336 (3rd Generation) |
| RT Cores | 84 (2nd Generation) |
| Interface | PCIe 4.0 x16 |
| Form Factor | Full-Height, Full-Length (FHFL) Dual-Slot |
| Thermal Design | Passive cooling (requires system airflow) |
| Max Thermal Design Power (TDP) | 300 W |
| NVLink Support | Yes (NVLink Bridge, up to 2 GPUs) |
| Display Outputs | 4x DisplayPort 1.4 |
| GPU Virtualization | NVIDIA vGPU supported |
| Multi-Instance GPU (MIG) | Not supported on A40 |
| ECC Memory Protection | Yes |
| API Support | CUDA, OpenCL, Vulkan, OpenGL, DirectX 12 Ultimate |
| Supported Server Platforms | HPE ProLiant DL380 Gen10 Plus, DL360 Gen10 Plus, and compatible HPE Apollo systems |
| Management Integration | HPE iLO compatible |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.