Brand: Intel | Category: GPUs
SKU: HLS-GAUDI2-PCIE | Part #: HLS-GAUDI2-PCIE | MPN: HLS-GAUDI2-PCIE
Contact for Pricing — Request a Quote
This Intel Gaudi 2 AI Accelerator PCIe Card operates at PCIe Gen 4 x16 with a 600 W thermal design power requirement—critical specifications that determine both your server chassis PCIe slot generation and power supply unit capacity during procurement and deployment. The HLS-GAUDI2-PCIE accelerator delivers 96 GB of HBM2e memory configured across 6 stacks, paired with 2.45 TB/s memory bandwidth to support demanding large-language model inference and training workloads. With 24 Tensor Processor Cores and a dual-core Matrix Multiplication Engine, this Intel processor is engineered for high-throughput compute tasks across distributed environments.
Built as a standard PCIe add-in card, the Gaudi 2 integrates directly into enterprise rack servers without custom modifications. The accelerator includes 24 integrated 100GbE ports supporting RoCE v2—21 dedicated to scale-out cluster connectivity and 3 reserved for host communication—enabling low-latency multi-GPU training and inference across data centers. Intel SynapseAI provides an open-source software stack including drivers, runtime, and graph compiler support for PyTorch and TensorFlow frameworks on Linux (Ubuntu, CentOS/RHEL) operating systems. This makes the Gaudi 2 an ideal choice for AI infrastructure teams seeking vendor-backed deep learning acceleration with native enterprise Linux support and standardized PCIe deployment.
For detailed specifications, compatibility verification, and enterprise procurement options, contact the Omnixon Global team to request a quotation.
| Brand | Intel |
| Category | GPUs |
| SKU | HLS-GAUDI2-PCIE |
| Part Number | HLS-GAUDI2-PCIE |
| Condition | New |
| Manufacturer Part Number | HLS-GAUDI2-PCIE |
| Product Family | Intel Gaudi 2 |
| Form Factor | PCIe Add-in Card |
| Memory Capacity | 96 GB HBM2e |
| Memory Configuration | 6 × HBM2e stacks |
| Memory Bandwidth | 2.45 TB/s |
| Tensor Processor Cores (TPCs) | 24 |
| Matrix Multiplication Engine (MME) | 1 (dual-core) |
| Host Interface | PCIe Gen 4 x16 |
| Integrated Network Ports | 24 × 100GbE (RoCE v2) |
| Scale-Out Ports | 21 × 100GbE RDMA |
| Host Ports | 3 × 100GbE |
| TDP (Thermal Design Power) | 600 W |
| Cooling | Active (requires server chassis airflow) |
| Supported Frameworks | PyTorch, TensorFlow (via Intel SynapseAI SDK) |
| Software Stack | Intel SynapseAI (open-source drivers, runtime, graph compiler) |
| Operating System Support | Linux (Ubuntu, CentOS/RHEL) |
| Server Compatibility | Standard PCIe Gen 4 enterprise rack servers |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.