Brand: Intel | Category: GPUs
SKU: HL-205 | Part #: HL-205 | MPN: HL-205
Contact for Pricing — Request a Quote
The Intel Habana Gaudi 1 HL-205 is an OAM (OCP Accelerator Module) form-factor deep learning training accelerator built on Habana Labs' first-generation Gaudi architecture. Designed specifically for large-scale AI training workloads in datacenter environments, the HL-205 integrates eight on-die GEMM engines alongside a matrix multiplication unit and a shared SRAM pool to deliver high-throughput tensor computations. The processor incorporates 10 x 100 GbE RoCE v2 ports directly on-chip, enabling scale-out training across multiple nodes without requiring a separate high-speed interconnect fabric, which simplifies cluster topology and reduces infrastructure complexity.
The Gaudi 1 architecture features a heterogeneous compute design combining dedicated Tensor Processing Cores (TPC) — fully programmable VLIW SIMD processors — with the Matrix Multiplication Engine (MME). The TPC cores support a broad range of data types including FP32, BF16, INT16, INT8, UINT8, INT4, and UINT4, providing flexibility for mixed-precision training workflows. The HL-205 carries 32 GB of HBM2 high-bandwidth memory with an aggregate memory bandwidth suited to feeding the high-throughput compute engines during large batch training operations. The OAM mechanical form factor aligns with OCP standards, facilitating integration into OAM-compatible baseboard platforms and high-density server chassis designed for AI infrastructure.
Targeted at enterprises building and operating large neural network training pipelines, the Gaudi 1 HL-205 supports popular deep learning frameworks including TensorFlow and PyTorch through Habana's SynapseAI software suite. The card's native scale-out networking, implemented via on-chip 100 GbE interfaces with RDMA over Converged Ethernet, allows multi-node training clusters to be constructed using standard Ethernet switching infrastructure. This approach is particularly relevant for organizations seeking to scale AI training capacity across UAE, GCC, EMEA, and APAC datacenter footprints without the operational overhead of proprietary interconnect ecosystems.
| Manufacturer | Intel |
| Brand Family | Intel Habana Gaudi 1 |
| Model | HL-205 |
| Form Factor | OAM (OCP Accelerator Module) |
| Architecture | Gaudi 1 |
| Compute Engines | 8 x Tensor Processing Cores (TPC) VLIW SIMD + Matrix Multiplication Engine (MME) |
| HBM Memory | 32 GB HBM2 |
| Supported Data Types | FP32, BF16, INT16, INT8, UINT8, INT4, UINT4 |
| On-Chip Scale-Out Networking | 10 x 100 GbE RoCE v2 (RDMA over Converged Ethernet) |
| Network Interface Standard | 100 Gigabit Ethernet, on-die integrated |
| Scale-Out Fabric | Standard Ethernet switching (no proprietary interconnect required) |
| Software Stack | Habana SynapseAI SDK |
| Framework Support | TensorFlow, PyTorch |
| OCP Compliance | OCP OAM v1.0 specification compliant |
| Target Workload | Deep learning training |
| Manufacturer Part Number | HL-205 |
Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.
| Brand | Intel |
| Category | GPUs |
| SKU | HL-205 |
| Part Number | HL-205 |
| Condition | New |
| Brand Family | Intel Habana Gaudi 1 |
| Model | HL-205 |
| Form Factor | OAM (OCP Accelerator Module) |
| Architecture | Gaudi 1 |
| Compute Engines | 8 x Tensor Processing Cores (TPC) VLIW SIMD + Matrix Multiplication Engine (MME) |
| HBM Memory | 32 GB HBM2 |
| Supported Data Types | FP32, BF16, INT16, INT8, UINT8, INT4, UINT4 |
| On-Chip Scale-Out Networking | 10 x 100 GbE RoCE v2 (RDMA over Converged Ethernet) |
| Network Interface Standard | 100 Gigabit Ethernet, on-die integrated |
| Scale-Out Fabric | Standard Ethernet switching (no proprietary interconnect required) |
| Software Stack | Habana SynapseAI SDK |
| Framework Support | TensorFlow, PyTorch |
| OCP Compliance | OCP OAM v1.0 specification compliant |
| Target Workload | Deep learning training |
| Manufacturer Part Number | HL-205 |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.