NVIDIA L40S 48GB GDDR6 PCIe Gen4 Datacenter GPU

NVIDIA L40S 48GB GDDR6 PCIe Gen4 Datacenter GPU

Brand: NVIDIA | Category: GPUs

SKU: 900-2G133-0020-000 | Part #: 900-2G133-0020-000 | MPN: 900-2G133-0020-000

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the NVIDIA L40S 48GB GDDR6 PCIe Gen4 Datacenter GPU

The NVIDIA L40S 48GB is a full-stack data center GPU built on the Ada Lovelace architecture, delivering exceptional compute density for AI training, inference, graphics rendering, and general-purpose HPC workloads. It features 18,176 CUDA cores, 568 Tensor Cores (4th generation), and 142 RT Cores (3rd generation), paired with 48GB of high-bandwidth GDDR6 memory across a 384-bit memory bus. Connecting via PCIe Gen 4 x16, the L40S is designed to slot into standard enterprise servers without requiring proprietary interconnect fabrics, making it highly deployable across existing data center infrastructure.

With a peak FP32 performance of 91.6 TFLOPS and FP16 Tensor Core performance reaching 362 TFLOPS (sparse), the L40S is engineered to handle demanding generative AI inference workloads, large language model (LLM) fine-tuning, and multi-modal AI pipelines. Its 48GB frame buffer supports large batch sizes and complex model architectures without memory bottlenecks, while NVIDIA's sparsity acceleration allows further throughput gains on compatible models. The GPU also supports INT8 at 733 TOPS (sparse) and FP8 at 1457 TOPS (sparse), enabling highly efficient quantized inference deployments at scale.

Beyond AI and HPC, the L40S retains full professional visualization capabilities inherited from the Ada Lovelace generation, including hardware-accelerated ray tracing and DLSS 3 support, making it suitable for enterprises requiring a unified platform across visual computing, simulation, and AI inference within the same server fleet. With a 350W TDP in a dual-slot form factor and support for NVLink Bridge for multi-GPU configurations, the L40S delivers enterprise-grade versatility for organizations standardizing on a single high-performance GPU SKU across multiple workload types.

Ideal for

  • Large language model (LLM) inference and fine-tuning in on-premises or private cloud data center environments
  • Generative AI application serving, including text-to-image and multimodal model deployments at enterprise scale
  • High-performance computing (HPC) workloads such as scientific simulation, computational fluid dynamics, and molecular dynamics
  • Professional 3D visualization, real-time ray tracing, and virtual workstation delivery via GPU-accelerated VDI platforms
  • AI-assisted video processing, transcoding, and computer vision pipelines requiring sustained throughput across large media datasets
  • Multi-tenant cloud GPU instances and enterprise AI platforms requiring dense compute per rack unit with broad server compatibility

Technical specifications

ManufacturerNVIDIA
Manufacturer Part Number900-2G133-0020-000
GPU ArchitectureAda Lovelace
CUDA Cores18,176
Tensor Cores568 (4th Generation)
RT Cores142 (3rd Generation)
Memory Size48 GB GDDR6
Memory Bus Width384-bit
Memory Bandwidth864 GB/s
Peak FP32 Performance91.6 TFLOPS
Peak FP16 Tensor Performance (Sparse)362.05 TFLOPS
Peak INT8 Performance (Sparse)733 TOPS
Peak FP8 Performance (Sparse)1457 TOPS
System InterfacePCIe Gen 4 x16
Form FactorDual-slot, full-height full-length (FHFL)
Thermal Design Power (TDP)350 W
Multi-GPU SupportNVLink Bridge (2-way)
Display Outputs4x DisplayPort 1.4
ECC Memory SupportYes
NVENC / NVDEC Engines2x NVENC, 2x NVDEC, 1x AV1 Encode
Operating System SupportLinux, Windows Server

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandNVIDIA
CategoryGPUs
SKU900-2G133-0020-000
Part Number900-2G133-0020-000
ConditionNew
Manufacturer Part Number900-2G133-0020-000
GPU ArchitectureAda Lovelace
CUDA Cores18,176
Tensor Cores568 (4th Generation)
RT Cores142 (3rd Generation)
Memory Size48 GB GDDR6
Memory Bus Width384-bit
Memory Bandwidth864 GB/s
Peak FP32 Performance91.6 TFLOPS
Peak FP16 Tensor Performance (Sparse)362.05 TFLOPS
Peak INT8 Performance (Sparse)733 TOPS
Peak FP8 Performance (Sparse)1457 TOPS
System InterfacePCIe Gen 4 x16
Form FactorDual-slot, full-height full-length (FHFL)
Thermal Design Power (TDP)350 W
Multi-GPU SupportNVLink Bridge (2-way)
Display Outputs4x DisplayPort 1.4
ECC Memory SupportYes
NVENC / NVDEC Engines2x NVENC, 2x NVDEC, 1x AV1 Encode
Operating System SupportLinux, Windows Server

Frequently Asked Questions about NVIDIA L40S 48GB GDDR6 PCIe Gen4 Datacenter GPU

What server platforms accept the NVIDIA L40S 48GB GDDR6 PCIe Gen4 Datacenter GPU?

Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.

How long is the lead time on AI GPUs?

Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.

Do you supply matched networking (Quantum InfiniBand / Spectrum-X)?

Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.

Can you help with NVIDIA AI Enterprise licensing?

Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.