Brand: HPE | Category: GPUs
SKU: P66210-B21 | Part #: P66210-B21 | MPN: P66210-B21
Contact for Pricing — Request a Quote
The HPE Intel Gaudi 3 OAM 96GB AI Accelerator (P66210-B21) is an Open Accelerator Module engineered with liquid cooling support for high-density AI inference and training workloads in enterprise data centers. Built on Intel Gaudi 3 architecture, this accelerator delivers 1835 TFLOPS of peak BF16 compute and 3670 TFLOPS of peak FP8 compute, powered by 96 GB of HBM2e memory with 3.7 TB/s bandwidth. The OAM form factor integrates seamlessly into HPE AI and data center servers, enabling rapid scaling without redesigning host infrastructure.
Networking performance is a critical strength: the module features 24x 100GbE RoCE v2 ports for ultra-low-latency, RDMA-enabled cluster communications essential in distributed training environments. The accelerator supports FP8, BF16, FP16, and FP32 precisions, accommodating diverse model requirements across PyTorch and TensorFlow frameworks via Intel SynapseAI SDK. With a 900W thermal design power and liquid cooling requirement, deployment planning for AI infrastructure teams must account for cooling provisioning alongside electrical capacity. HPE's OAM specification ensures vendor-neutral interconnect compatibility, reducing lock-in risk for large-scale deployments.
Contact Omnixon Global to discuss HPE Intel Gaudi 3 OAM 96GB AI Accelerator availability, configuration guidance, and enterprise licensing options.
| Brand | HPE |
| Category | GPUs |
| SKU | P66210-B21 |
| Part Number | P66210-B21 |
| Condition | New |
| Manufacturer Part Number | P66210-B21 |
| Product Name | HPE Intel Gaudi 3 OAM 96GB AI Accelerator |
| Accelerator Architecture | Intel Gaudi 3 |
| Form Factor | OAM (Open Accelerator Module) |
| HBM Capacity | 96 GB |
| Memory Type | HBM2e |
| Memory Bandwidth | 3.7 TB/s |
| Peak BF16 Compute | 1835 TFLOPS |
| Peak FP8 Compute | 3670 TFLOPS |
| Integrated Network Ports | 24x 100GbE RoCE v2 |
| Network Protocol | RoCE v2 (RDMA over Converged Ethernet) |
| Supported Precisions | FP8, BF16, FP16, FP32 |
| AI Framework Support | PyTorch, TensorFlow (via Intel SynapseAI SDK) |
| Interconnect Standard | Open Accelerator Module (OAM) specification |
| TDP | 900W |
| Cooling | Liquid cooling required (OAM module) |
| Target Platform | HPE AI and data center servers supporting OAM form factor |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.