Omnixon Global
PNY NVIDIA L40S 48GB PCIe GPU Accelerator

PNY NVIDIA L40S 48GB PCIe GPU Accelerator

Brand: NVIDIA | Category: GPUs

SKU: TCSL40S48GB-1-KIT | Part #: TCSL40S48GB-1-KIT | MPN: TCSL40S48GB-1-KIT

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the PNY NVIDIA L40S 48GB PCIe GPU Accelerator

The NVIDIA L40S is a full-height, dual-slot PCIe GPU accelerator purpose-built for enterprise data center inference, generative AI deployment, and professional visualization workloads. Based on NVIDIA's Ada Lovelace architecture, the L40S delivers 568 TFLOPS of FP32 performance and 142 TFLOPS of sparsity-accelerated compute, with 48GB of GDDR6 memory and 384-bit memory bandwidth optimized for large language model serving and batch inference at scale.

The L40S features 18,176 CUDA cores, enabling efficient multi-instance GPU (MIG) partitioning for workload consolidation and resource isolation in shared infrastructure. Native support for tensor operations, dynamic shape inference, and mixed-precision compute (FP32, TF32, FP16, BF16, INT8) accelerates transformer-based models and retrieval-augmented generation (RAG) pipelines across production deployments.

Designed for 295W peak power consumption in standard ATX and modular data center chassis, the L40S integrates NVIDIA NVLink Switch fabric compatibility for multi-GPU coherence and supports industry-standard PCIe Gen 4 x16 connectivity. Extended memory capacity and compute density make it the preferred GPU for long-context LLM inference, token generation, and concurrent multi-model serving in hyperscaler and enterprise AI infrastructure.

Ideal for

  • Large language model (LLM) inference serving and token generation at scale
  • Retrieval-augmented generation (RAG) and semantic search acceleration
  • Multi-tenant AI model serving with workload isolation via MIG partitioning
  • Enterprise generative AI applications (chatbots, content generation, code assistance)
  • Batch inference and offline AI processing for analytics and recommendation systems
  • Professional visualization and ray-tracing rendering in data center environments

Technical specifications

ManufacturerNVIDIA
Product LineL-Series
ModelL40S
ArchitectureAda Lovelace
Memory48GB GDDR6
Memory Bandwidth864 GB/s
CUDA Cores18176
FP32 Performance568 TFLOPS
Sparsity-Accelerated Compute142 TFLOPS
Tensor Performance (TF32)568 TFLOPS
Power Consumption295W
Form FactorDual-slot, full-height
InterfacePCIe Gen 4 x16
Multi-Instance GPU (MIG)Supported
Max Memory Allocation per Instance12GB or 24GB partitions
NVLink SupportCompatible with NVIDIA NVLink Switch
Max Concurrent Threads per Block1024
ECC MemoryYes

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandNVIDIA
CategoryGPUs
SKUTCSL40S48GB-1-KIT
Part NumberTCSL40S48GB-1-KIT
ConditionNew
Product LineL-Series
ModelL40S
ArchitectureAda Lovelace
Memory48GB GDDR6
Memory Bandwidth864 GB/s
CUDA Cores18176
FP32 Performance568 TFLOPS
Sparsity-Accelerated Compute142 TFLOPS
Tensor Performance (TF32)568 TFLOPS
Power Consumption295W
Form FactorDual-slot, full-height
InterfacePCIe Gen 4 x16
Multi-Instance GPU (MIG)Supported
Max Memory Allocation per Instance12GB or 24GB partitions
NVLink SupportCompatible with NVIDIA NVLink Switch
Max Concurrent Threads per Block1024
ECC MemoryYes