Intel Gaudi 3 Inference Reference Design 4-Card PCIe

Intel Gaudi 3 Inference Reference Design 4-Card PCIe

Brand: Intel | Category: GPUs

SKU: HLRD-325B-4 | Part #: HLRD-325B-4 | MPN: HLRD-325B-4

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the Intel Gaudi 3 Inference Reference Design 4-Card PCIe

The Intel Gaudi 3 Inference Reference Design 4-Card PCIe (HLRD-325B-4) is a validated, production-ready inference platform integrating four Intel Gaudi 3 AI accelerators in a PCIe form-factor configuration. Built on Intel's third-generation Gaudi architecture, Gaudi 3 is fabricated on a 5nm process node and delivers substantial improvements in compute throughput and memory bandwidth over its predecessor. Each Gaudi 3 accelerator features 64 Tensor Processor Cores (TPCs) and a Matrix Multiplication Engine (MME) optimized for both FP8 and BF16 precision, enabling high-throughput execution of large language model (LLM) inference, generative AI, and deep learning recommendation workloads at enterprise scale.

The reference design leverages Gaudi 3's integrated RDMA-capable 21-port Ethernet fabric — 24 × 100GbE ports per card — to enable direct card-to-card communication without requiring an external network switch in scale-up configurations. Each accelerator is equipped with 128 GB of HBM2e memory, providing the high-capacity, high-bandwidth memory footprint required for hosting large generative models in their entirety. The 4-card PCIe assembly is validated as a complete subsystem, reducing integration risk for datacenter operators deploying inference infrastructure based on standard PCIe server platforms.

Designed for enterprise datacenters and cloud service providers, the HLRD-325B-4 is supported by Intel's open software ecosystem including the Intel Gaudi Software Suite and compatibility with industry-standard frameworks such as PyTorch and Hugging Face Transformers via the Optimum Habana integration. This reference design is particularly well-suited for organizations seeking a scalable, standards-based PCIe inference solution for production AI services without proprietary interconnect lock-in.

Ideal for

  • Large language model (LLM) inference serving for enterprise generative AI applications, including multi-billion parameter transformer models
  • High-throughput text generation and summarization pipelines in financial services, legal, and media organizations
  • Retrieval-augmented generation (RAG) inference backends requiring large on-accelerator memory capacity to host both retriever and generator models simultaneously
  • Deep learning recommendation model (DLRM) inference for e-commerce, digital advertising, and content personalization platforms
  • Multi-modal AI inference workloads combining vision and language models in healthcare imaging, manufacturing quality control, and document processing
  • Enterprise AI inference infrastructure consolidation in datacenter environments standardized on PCIe server architectures, avoiding proprietary fabric requirements

Technical specifications

ManufacturerIntel
Manufacturer Part NumberHLRD-325B-4
Product FamilyIntel Gaudi 3
Form FactorPCIe Reference Design (4-Card Assembly)
Number of Accelerators4 × Intel Gaudi 3
Process Node5nm
Tensor Processor Cores (TPC) per Card64
HBM Memory per Card128 GB HBM2e
Total HBM Memory (4-Card)512 GB HBM2e
Supported PrecisionsFP8, BF16, FP16, FP32, INT16, INT8
On-Card Networking24 × 100GbE RDMA-capable ports per card
Host InterfacePCIe Gen 4
Scale-Up InterconnectIntegrated Ethernet-based RDMA fabric (no external switch required for scale-up)
Software Framework SupportPyTorch, Hugging Face Optimum Habana, Intel Gaudi Software Suite
Operating System SupportLinux (Ubuntu, RHEL)
Target WorkloadAI Inference (LLM, Generative AI, DLRM)
Configuration TypeValidated Inference Reference Design
CoolingActive (air-cooled)

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandIntel
CategoryGPUs
SKUHLRD-325B-4
Part NumberHLRD-325B-4
ConditionNew
Manufacturer Part NumberHLRD-325B-4
Product FamilyIntel Gaudi 3
Form FactorPCIe Reference Design (4-Card Assembly)
Number of Accelerators4 × Intel Gaudi 3
Process Node5nm
Tensor Processor Cores (TPC) per Card64
HBM Memory per Card128 GB HBM2e
Total HBM Memory (4-Card)512 GB HBM2e
Supported PrecisionsFP8, BF16, FP16, FP32, INT16, INT8
On-Card Networking24 × 100GbE RDMA-capable ports per card
Host InterfacePCIe Gen 4
Scale-Up InterconnectIntegrated Ethernet-based RDMA fabric (no external switch required for scale-up)
Software Framework SupportPyTorch, Hugging Face Optimum Habana, Intel Gaudi Software Suite
Operating System SupportLinux (Ubuntu, RHEL)
Target WorkloadAI Inference (LLM, Generative AI, DLRM)
Configuration TypeValidated Inference Reference Design
CoolingActive (air-cooled)

Frequently Asked Questions about Intel Gaudi 3 Inference Reference Design 4-Card PCIe

What server platforms accept the Intel Gaudi 3 Inference Reference Design 4-Card PCIe?

Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.

How long is the lead time on AI GPUs?

Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.

Do you supply matched networking (Quantum InfiniBand / Spectrum-X)?

Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.

Can you help with NVIDIA AI Enterprise licensing?

Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.