Brand: HPE | Category: GPUs
SKU: P53866-B21 | Part #: P53866-B21 | MPN: P53866-B21
Contact for Pricing — Request a Quote
The HPE NVIDIA L40S 48GB PCIe GPU Accelerator delivers 48 GB of GDDR6 ECC memory with 864 GB/s bandwidth, providing substantial compute capacity for demanding enterprise workloads. Built on NVIDIA's Ada Lovelace architecture, this dual-slot PCIe Gen 4 x16 accelerator combines high-performance tensor operations with efficient single-precision compute, making it ideal for AI inference, data analytics, and professional visualization tasks in virtualized environments.
With 18,176 CUDA cores and 4th Generation Tensor Cores, the HPE NVIDIA L40S (part number P53866-B21) delivers impressive throughput across multiple precision levels. The 3rd Generation RT Cores enable real-time ray tracing acceleration, while the passive cooling design—dependent on server airflow—simplifies deployment in standard HPE server chassis. This accelerator is engineered for organizations seeking to enhance inference performance and data processing capabilities without the complexity of multi-GPU clustering via NVLink. IT infrastructure teams and AI platform architects will find this solution particularly valuable for scaling virtualized compute environments. Contact Omnixon Global today to request a quote.
| Brand | HPE |
| Category | GPUs |
| SKU | P53866-B21 |
| Part Number | P53866-B21 |
| Condition | New |
| Manufacturer Part Number | P53866-B21 |
| Product Name | HPE NVIDIA L40S 48GB PCIe GPU Accelerator |
| GPU Architecture | NVIDIA Ada Lovelace |
| CUDA Cores | 18176 |
| GPU Memory | 48 GB GDDR6 ECC |
| Memory Bandwidth | 864 GB/s |
| Memory Interface | 384-bit |
| FP32 Performance | 91.6 TFLOPS |
| TF32 Tensor Core Performance | 183 TFLOPS (sparsity: 366 TFLOPS) |
| FP16 Tensor Core Performance | 362.05 TFLOPS (sparsity: 724.1 TFLOPS) |
| INT8 Tensor Core Performance | 724.1 TOPS (sparsity: 1457.9 TOPS) |
| RT Cores (Generation) | 3rd Generation |
| Tensor Cores (Generation) | 4th Generation |
| NVLink Support | None (PCIe only) |
| Form Factor | Dual-slot PCIe |
| Interface | PCIe Gen 4 x16 |
| Thermal Design Power (TDP) | 350 W |
| Cooling | Passive (server airflow dependent) |
| Display Outputs | None (compute and visualization via virtualization; no physical display connectors) |
| Virtualization Support | NVIDIA vGPU (vWS, vPC, vCS profiles supported) |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.