Brand: NVIDIA | Category: GPUs
SKU: 900-2G133-0020-000 | Part #: 900-2G133-0020-000 | MPN: 900-2G133-0020-000
Contact for Pricing — Request a Quote
The NVIDIA L40S 48GB is a full-stack data center GPU built on the Ada Lovelace architecture, delivering exceptional compute density for AI training, inference, graphics rendering, and general-purpose HPC workloads. It features 18,176 CUDA cores, 568 Tensor Cores (4th generation), and 142 RT Cores (3rd generation), paired with 48GB of high-bandwidth GDDR6 memory across a 384-bit memory bus. Connecting via PCIe Gen 4 x16, the L40S is designed to slot into standard enterprise servers without requiring proprietary interconnect fabrics, making it highly deployable across existing data center infrastructure.
With a peak FP32 performance of 91.6 TFLOPS and FP16 Tensor Core performance reaching 362 TFLOPS (sparse), the L40S is engineered to handle demanding generative AI inference workloads, large language model (LLM) fine-tuning, and multi-modal AI pipelines. Its 48GB frame buffer supports large batch sizes and complex model architectures without memory bottlenecks, while NVIDIA's sparsity acceleration allows further throughput gains on compatible models. The GPU also supports INT8 at 733 TOPS (sparse) and FP8 at 1457 TOPS (sparse), enabling highly efficient quantized inference deployments at scale.
Beyond AI and HPC, the L40S retains full professional visualization capabilities inherited from the Ada Lovelace generation, including hardware-accelerated ray tracing and DLSS 3 support, making it suitable for enterprises requiring a unified platform across visual computing, simulation, and AI inference within the same server fleet. With a 350W TDP in a dual-slot form factor and support for NVLink Bridge for multi-GPU configurations, the L40S delivers enterprise-grade versatility for organizations standardizing on a single high-performance GPU SKU across multiple workload types.
| Manufacturer | NVIDIA |
| Manufacturer Part Number | 900-2G133-0020-000 |
| GPU Architecture | Ada Lovelace |
| CUDA Cores | 18,176 |
| Tensor Cores | 568 (4th Generation) |
| RT Cores | 142 (3rd Generation) |
| Memory Size | 48 GB GDDR6 |
| Memory Bus Width | 384-bit |
| Memory Bandwidth | 864 GB/s |
| Peak FP32 Performance | 91.6 TFLOPS |
| Peak FP16 Tensor Performance (Sparse) | 362.05 TFLOPS |
| Peak INT8 Performance (Sparse) | 733 TOPS |
| Peak FP8 Performance (Sparse) | 1457 TOPS |
| System Interface | PCIe Gen 4 x16 |
| Form Factor | Dual-slot, full-height full-length (FHFL) |
| Thermal Design Power (TDP) | 350 W |
| Multi-GPU Support | NVLink Bridge (2-way) |
| Display Outputs | 4x DisplayPort 1.4 |
| ECC Memory Support | Yes |
| NVENC / NVDEC Engines | 2x NVENC, 2x NVDEC, 1x AV1 Encode |
| Operating System Support | Linux, Windows Server |
Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.
| Brand | NVIDIA |
| Category | GPUs |
| SKU | 900-2G133-0020-000 |
| Part Number | 900-2G133-0020-000 |
| Condition | New |
| Manufacturer Part Number | 900-2G133-0020-000 |
| GPU Architecture | Ada Lovelace |
| CUDA Cores | 18,176 |
| Tensor Cores | 568 (4th Generation) |
| RT Cores | 142 (3rd Generation) |
| Memory Size | 48 GB GDDR6 |
| Memory Bus Width | 384-bit |
| Memory Bandwidth | 864 GB/s |
| Peak FP32 Performance | 91.6 TFLOPS |
| Peak FP16 Tensor Performance (Sparse) | 362.05 TFLOPS |
| Peak INT8 Performance (Sparse) | 733 TOPS |
| Peak FP8 Performance (Sparse) | 1457 TOPS |
| System Interface | PCIe Gen 4 x16 |
| Form Factor | Dual-slot, full-height full-length (FHFL) |
| Thermal Design Power (TDP) | 350 W |
| Multi-GPU Support | NVLink Bridge (2-way) |
| Display Outputs | 4x DisplayPort 1.4 |
| ECC Memory Support | Yes |
| NVENC / NVDEC Engines | 2x NVENC, 2x NVDEC, 1x AV1 Encode |
| Operating System Support | Linux, Windows Server |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.