Intel Gaudi 3 PCIe AI Accelerator (HL-388)

Intel Gaudi 3 PCIe AI Accelerator (HL-388)

Brand: Intel | Category: GPUs

SKU: INTE-HL388 | Part #: HL-388 | MPN: HL-388

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the Intel Gaudi 3 PCIe AI Accelerator (HL-388)

The Intel Gaudi 3 HL-388 is a PCIe Gen 5 x16 add-in-card that places the full Gaudi 3 AI compute die into a standard double-slot, full-height/full-length (FHFL) form factor, enabling enterprises to retrofit existing rack-mount servers with purpose-built AI acceleration without requiring proprietary OAM baseboards or specialized infrastructure. At its core sits Intel's second-generation Gaudi 3 ASIC built on TSMC 5nm process technology, featuring 64 Tensor Processor Cores (TPCs) and a Matrix Multiplication Engine (MME) capable of delivering up to 1,835 TFLOPS in BF16 and 3,670 TOPS in INT8, making it a compelling option for large language model inference and fine-tuning workloads at scale.

The HL-388 is equipped with 96 GB of HBM2e memory delivering 3.7 TB/s of memory bandwidth, which supports running large parameter models directly on a single card while maintaining high throughput. The card integrates 24 integrated 200 Gb/s RDMA-capable network ports accessible via external QSFP-DD connectors, enabling low-latency, high-bandwidth scale-out across multi-node clusters through the Gaudi 3 native scale-out fabric — eliminating the need for third-party InfiniBand or RoCE network switches in many topologies. Thermal design is handled through an active cooling solution rated at a 600W TDP, with power delivery via a standard PCIe slot plus supplemental PCIe power connectors.

From a software ecosystem perspective, the HL-388 is fully supported by Intel's SynapseAI SDK, which provides optimized deep learning framework integration with PyTorch and TensorFlow through the Habana deep learning compiler. The card is compatible with Hugging Face Optimum Habana, enabling straightforward deployment of transformer-based models. The PCIe form factor specifically targets enterprises seeking to expand AI inference capacity in standard 2U and 4U rack servers from leading brands, avoiding the capital expenditure associated with full OAM-based Gaudi 3 node deployments while still accessing the same underlying silicon performance.

Ideal for

  • Large language model (LLM) inference serving for enterprise generative AI applications, running models such as Llama 3, Mixtral, and similar architectures at production throughput levels
  • LLM fine-tuning and domain adaptation using parameter-efficient techniques (LoRA, QLoRA) on proprietary enterprise datasets without dedicated AI server infrastructure
  • Retrieval-Augmented Generation (RAG) pipeline acceleration combining embedding model inference with real-time document retrieval for enterprise knowledge management platforms
  • Multi-modal AI model inference including vision-language models and speech recognition systems deployed within existing enterprise data center servers
  • AI-assisted code generation and developer productivity tooling hosted on-premises to satisfy data residency and IP security requirements
  • Expanding AI compute capacity in co-location facilities where OAM-based specialized AI servers are not available or permitted, using standard 2U/4U server chassis already under lease

Technical specifications

ManufacturerIntel
Product LineIntel Gaudi 3
Model / SKUHL-388
Form FactorPCIe Add-In Card (FHFL, Double-Slot)
Host InterfacePCIe Gen 5 x16
Process NodeTSMC 5nm
Tensor Processor Cores (TPCs)64
BF16 Compute Performance1,835 TFLOPS
FP8 Compute Performance3,670 TFLOPS
INT8 Compute Performance3,670 TOPS
Memory TypeHBM2e
Memory Capacity96 GB
Memory Bandwidth3.7 TB/s
Scale-Out Networking24 × 200 Gb/s RDMA ports (integrated), via QSFP-DD connectors
Scale-Out Fabric BandwidthUp to 4.8 Tb/s aggregate bidirectional
Thermal Design Power (TDP)600 W
CoolingActive (forced-air, onboard fans + server airflow)
Power ConnectorsPCIe slot + supplemental PCIe power connectors
Supported FrameworksPyTorch, TensorFlow (via SynapseAI SDK and Optimum Habana)
Operating System SupportUbuntu 22.04 LTS, Red Hat Enterprise Linux 8/9
Virtualization SupportSR-IOV capable
Management InterfacePCIe management, Intel Gaudi driver stack, Prometheus-compatible telemetry
ComplianceRoHS, CE, FCC

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandIntel
CategoryGPUs
SKUINTE-HL388
Part NumberHL-388
ConditionNew
Product LineIntel Gaudi 3
Model / SKUHL-388
Form FactorPCIe Add-In Card (FHFL, Double-Slot)
Host InterfacePCIe Gen 5 x16
Process NodeTSMC 5nm
Tensor Processor Cores (TPCs)64
BF16 Compute Performance1,835 TFLOPS
FP8 Compute Performance3,670 TFLOPS
INT8 Compute Performance3,670 TOPS
Memory TypeHBM2e
Memory Capacity96 GB
Memory Bandwidth3.7 TB/s
Scale-Out Networking24 × 200 Gb/s RDMA ports (integrated), via QSFP-DD connectors
Scale-Out Fabric BandwidthUp to 4.8 Tb/s aggregate bidirectional
Thermal Design Power (TDP)600 W
CoolingActive (forced-air, onboard fans + server airflow)
Power ConnectorsPCIe slot + supplemental PCIe power connectors
Supported FrameworksPyTorch, TensorFlow (via SynapseAI SDK and Optimum Habana)
Operating System SupportUbuntu 22.04 LTS, Red Hat Enterprise Linux 8/9
Virtualization SupportSR-IOV capable
Management InterfacePCIe management, Intel Gaudi driver stack, Prometheus-compatible telemetry
ComplianceRoHS, CE, FCC

Frequently Asked Questions about Intel Gaudi 3 PCIe AI Accelerator (HL-388)

What does the Intel Gaudi 3 PCIe AI Accelerator (HL-388) do?

The Intel Gaudi 3 PCIe AI Accelerator (HL-388) accelerates AI/ML training, inference, scientific HPC, and virtualization (vGPU) workloads. Typical deployments include LLM training clusters, computer-vision pipelines, financial risk modeling, and rendering farms.

What does the Intel Gaudi 3 PCIe AI Accelerator (HL-388) do?

The Intel Gaudi 3 PCIe AI Accelerator (HL-388) accelerates AI/ML training, inference, scientific HPC, and virtualization (vGPU) workloads. Typical deployments include LLM training clusters, computer-vision pipelines, financial risk modeling, and rendering farms.

What does the Intel Gaudi 3 PCIe AI Accelerator (HL-388) do?

The Intel Gaudi 3 PCIe AI Accelerator (HL-388) accelerates AI/ML training, inference, scientific HPC, and virtualization (vGPU) workloads. Typical deployments include LLM training clusters, computer-vision pipelines, financial risk modeling, and rendering farms.

What are the headline specs of the Intel Gaudi 3 PCIe AI Accelerator (HL-388)?

Key specifications for the Intel Gaudi 3 PCIe AI Accelerator (HL-388): new condition; manufacturer Intel; product line Intel Gaudi 3; model / sku HL-388; form factor PCIe Add-In Card (FHFL, Double-Slot); host interface PCIe Gen 5 x16; process node TSMC 5nm. Manufacturer part number HL-388. For the full datasheet with electrical, environmental, and compliance details, contact our pre-sales engineering team.

What are the headline specs of the Intel Gaudi 3 PCIe AI Accelerator (HL-388)?

Key specifications for the Intel Gaudi 3 PCIe AI Accelerator (HL-388): new condition; manufacturer Intel; product line Intel Gaudi 3; model / sku HL-388; form factor PCIe Add-In Card (FHFL, Double-Slot); host interface PCIe Gen 5 x16; process node TSMC 5nm. Manufacturer part number HL-388. For the full datasheet with electrical, environmental, and compliance details, contact our pre-sales engineering team.