Brand: Intel | Category: GPUs
SKU: HL-100 | Part #: HL-100 | MPN: HL-100
Contact for Pricing — Request a Quote
The Intel Habana Goya HL-100 is a dedicated inference accelerator designed for high-throughput deep learning inference workloads in datacenter and enterprise AI environments. Built on Habana Labs' first-generation Goya architecture, the HL-100 delivers a purpose-built execution engine optimized for neural network inference rather than general-purpose compute, enabling efficient processing of trained models across image classification, natural language processing, and recommendation system workloads. The card connects via a standard PCIe interface, making it compatible with a broad range of x86-based server platforms without requiring proprietary interconnects or specialized host infrastructure.
The Goya HL-100 integrates Habana's Tensor Processor Core (TPC), a VLIW SIMD processor specifically engineered for tensor operations central to deep learning inference. It supports both FP32 and INT8 data types, allowing operators to deploy full-precision models or leverage INT8 quantization to significantly increase throughput for latency-sensitive production inference pipelines. The device includes 32 GB of on-board HBM2 memory, providing substantial capacity for large model weights and activations without frequent host memory transfers, which is critical for batch inference serving scenarios.
As part of Intel's Habana portfolio, the HL-100 is supported by the SynapseAI software framework, which provides graph compilation, runtime management, and integration with standard deep learning frameworks including TensorFlow and PyTorch via ONNX import. This enables enterprise ML engineering teams to deploy existing trained models with minimal rework. The card is suited for deployment in AI inference servers, cloud service provider nodes, and on-premises AI inferencing appliances serving high-demand production traffic across industries including financial services, healthcare imaging, and large-scale internet services.
| Manufacturer | Intel |
| Brand | Habana Labs (Intel) |
| Model | Goya HL-100 |
| Manufacturer Part Number | HL-100 |
| Product Type | AI Inference Accelerator |
| Architecture | Habana Goya (Tensor Processor Core / TPC) |
| Host Interface | PCIe 3.0 x16 |
| Memory Capacity | 32 GB HBM2 |
| Supported Precisions | FP32, INT8 |
| Peak INT8 Throughput | 15,000 images/sec (ResNet-50, reported) |
| Peak FP32 Compute | ~700 GFLOPS |
| TDP (Thermal Design Power) | 200 W |
| Power Connector | Auxiliary PCIe power required |
| Form Factor | Full-height, full-length (FHFL) PCIe card |
| Cooling | Active (onboard fan) |
| Software Framework | SynapseAI SDK |
| Framework Support | TensorFlow, PyTorch (via ONNX) |
| Operating System Support | Linux (Ubuntu, CentOS) |
| Target Environment | Datacenter / Enterprise Server |
Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.
| Brand | Intel |
| Category | GPUs |
| SKU | HL-100 |
| Part Number | HL-100 |
| Condition | New |
| Model | Goya HL-100 |
| Manufacturer Part Number | HL-100 |
| Product Type | AI Inference Accelerator |
| Architecture | Habana Goya (Tensor Processor Core / TPC) |
| Host Interface | PCIe 3.0 x16 |
| Memory Capacity | 32 GB HBM2 |
| Supported Precisions | FP32, INT8 |
| Peak INT8 Throughput | 15,000 images/sec (ResNet-50, reported) |
| Peak FP32 Compute | ~700 GFLOPS |
| TDP (Thermal Design Power) | 200 W |
| Power Connector | Auxiliary PCIe power required |
| Form Factor | Full-height, full-length (FHFL) PCIe card |
| Cooling | Active (onboard fan) |
| Software Framework | SynapseAI SDK |
| Framework Support | TensorFlow, PyTorch (via ONNX) |
| Operating System Support | Linux (Ubuntu, CentOS) |
| Target Environment | Datacenter / Enterprise Server |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.