Brand: HPE | Category: GPUs
SKU: P66211-B21 | Part #: P66211-B21 | MPN: P66211-B21
Contact for Pricing — Request a Quote
The HPE Intel Gaudi 3 PCIe 96GB AI Accelerator delivers 96 GB of HBM2e memory, providing the high-bandwidth memory capacity essential for demanding AI workloads. This PCIe Gen 5 add-in card integrates Intel Gaudi 3 architecture to accelerate large language model training and inference, generative AI applications, computer vision, and distributed deep learning at enterprise scale.
Built for HPE ProLiant servers and HPE Cray systems with PCIe Gen 5 slots, the Gaudi 3 accelerator (part number P66211-B21) combines 24× 100GbE RDMA network interfaces with support for PyTorch and TensorFlow via the Intel Gaudi SynapseAI SDK. The card supports FP32, BF16, and FP8 precision formats, enabling flexible model optimization across inference and training scenarios. Native RoCE v2 scale-out networking facilitates seamless multi-accelerator deployments for distributed deep learning. Linux operating system support and HPE iLO ecosystem integration ensure straightforward deployment within existing data center infrastructure. Active cooling with adequate server airflow per HPE platform specifications maintains thermal performance under sustained workloads. This accelerator is purpose-built for AI infrastructure teams seeking to deploy high-performance, energy-efficient acceleration without proprietary software lock-in.
To evaluate this accelerator for your infrastructure requirements, contact the Omnixon Global team for a customized request for quotation.
| Brand | HPE |
| Category | GPUs |
| SKU | P66211-B21 |
| Part Number | P66211-B21 |
| Condition | New |
| Manufacturer Part Number | P66211-B21 |
| Product Name | HPE Intel Gaudi 3 PCIe 96GB AI Accelerator |
| AI Accelerator Architecture | Intel Gaudi 3 |
| Form Factor | PCIe Add-in Card |
| PCIe Interface | PCIe Gen 5 |
| Memory Capacity | 96 GB HBM2e |
| Memory Type | HBM2e (High Bandwidth Memory) |
| Supported Precision Formats | FP32, BF16, FP8 |
| On-Chip Network Interfaces | 24× 100GbE RDMA (RoCE) |
| Scale-Out Networking | RoCE v2 (RDMA over Converged Ethernet) |
| Supported AI Frameworks | PyTorch, TensorFlow (via Intel Gaudi SynapseAI SDK) |
| Thermal Design | Active cooling (requires adequate server airflow per HPE platform specifications) |
| Compatible Platforms | HPE ProLiant servers, HPE Cray systems with PCIe Gen 5 slots |
| Operating System Support | Linux (validated distributions per Intel Gaudi Software release notes) |
| Management Integration | HPE iLO ecosystem compatible |
| Target Workloads | LLM training and inference, generative AI, computer vision, distributed deep learning |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.