Quanta Computer Intel Gaudi 3 8-OAM HGX-Style Server QCT-G3

Quanta Computer Intel Gaudi 3 8-OAM HGX-Style Server QCT-G3

Brand: Intel | Category: GPUs

SKU: QCT-G3-8OAM | Part #: QCT-G3-8OAM | MPN: QCT-G3-8OAM

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the Quanta Computer Intel Gaudi 3 8-OAM HGX-Style Server QCT-G3

The Quanta Computer Intel Gaudi 3 8-OAM HGX-Style Server (QCT-G3-8OAM) is a purpose-built AI training and inference platform integrating eight Intel Gaudi 3 accelerator modules in an OAM (Open Accelerator Module) form factor within an industry-standard HGX-compatible server chassis. Each Gaudi 3 accelerator is built on a 5nm process node and delivers substantial improvements in compute density and memory bandwidth over its predecessor, featuring 64 Tensor Processor Cores (TPCs) and Matrix Multiplication Engines (MMEs) optimized for mixed-precision deep learning operations including BF16, FP8, and FP32. The eight accelerators are interconnected via Intel's high-bandwidth, low-latency scale-up fabric using 24 ports of 200Gbps Ethernet per OAM, enabling direct all-to-all communication without requiring external switch infrastructure for within-node collective operations.

The QCT-G3-8OAM platform is engineered to address the demands of large-scale generative AI model training, large language model (LLM) fine-tuning, and high-throughput inference serving. Each Gaudi 3 OAM carries 128GB of HBM2e memory, yielding an aggregate of 1TB of HBM capacity across the full eight-accelerator configuration, alongside aggregate memory bandwidth well suited to the memory-bound kernels common in transformer-based workloads. The server integrates dual high-core-count Intel Xeon Scalable processors to handle host-side data preprocessing, orchestration, and I/O, with PCIe Gen 5 connectivity between host CPUs and the accelerator complex.

Designed for enterprise datacenter deployment, the QCT-G3 platform supports standard data center power and cooling infrastructure, with the OAM modules utilizing direct liquid cooling or high-airflow air cooling depending on deployment configuration. The system is compatible with Intel's Gaudi software ecosystem, including the Intel Gaudi Software Suite, Optimum Habana integration for Hugging Face workloads, and support for PyTorch and TensorFlow via SynapseAI. Omnixon Global supplies this platform to enterprise and hyperscale customers across the UAE, GCC, EMEA, and APAC regions.

Ideal for

  • Large language model (LLM) pre-training and fine-tuning at scale, supporting multi-billion parameter transformer architectures such as LLaMA, Falcon, and GPT-class models
  • High-throughput generative AI inference serving for enterprise applications requiring low-latency responses across concurrent user sessions
  • Computer vision and multimodal model training for industrial inspection, medical imaging analysis, and autonomous systems development
  • Retrieval-augmented generation (RAG) pipeline acceleration combining dense embedding computation with transformer inference in enterprise AI platforms
  • Scientific computing and simulation workloads in HPC environments that benefit from high-bandwidth memory and mixed-precision tensor operations
  • Enterprise AI model experimentation and MLOps pipelines requiring reproducible, high-throughput training runs with broad deep learning framework compatibility

Technical specifications

ManufacturerIntel
Manufacturer Part NumberQCT-G3-8OAM
Accelerator ModelIntel Gaudi 3
Number of Accelerators8 x OAM modules
Accelerator Process Node5nm
Tensor Processor Cores (TPC) per OAM64
HBM Capacity per OAM128 GB HBM2e
Total Aggregate HBM Capacity1 TB (8 x 128 GB)
Scale-Up Interconnect per OAM24 x 200 Gbps Ethernet ports (built-in)
Scale-Out Network Ports per OAM2 x 400 Gbps Ethernet OSFP ports
Host CPUDual Intel Xeon Scalable Processors (5th Gen)
Host-to-Accelerator InterfacePCIe Gen 5
Supported PrecisionsFP32, BF16, FP16, FP8
Chassis Form FactorHGX-style OAM server
Software EcosystemIntel Gaudi SynapseAI SDK, PyTorch, TensorFlow, Optimum Habana
Framework SupportPyTorch, TensorFlow, Hugging Face Transformers (via Optimum Habana)
Cooling SupportAir cooling / Direct Liquid Cooling (DLC) depending on configuration
Target WorkloadsAI training, LLM fine-tuning, generative AI inference, HPC

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandIntel
CategoryGPUs
SKUQCT-G3-8OAM
Part NumberQCT-G3-8OAM
ConditionNew
Manufacturer Part NumberQCT-G3-8OAM
Accelerator ModelIntel Gaudi 3
Number of Accelerators8 x OAM modules
Accelerator Process Node5nm
Tensor Processor Cores (TPC) per OAM64
HBM Capacity per OAM128 GB HBM2e
Total Aggregate HBM Capacity1 TB (8 x 128 GB)
Scale-Up Interconnect per OAM24 x 200 Gbps Ethernet ports (built-in)
Scale-Out Network Ports per OAM2 x 400 Gbps Ethernet OSFP ports
Host CPUDual Intel Xeon Scalable Processors (5th Gen)
Host-to-Accelerator InterfacePCIe Gen 5
Supported PrecisionsFP32, BF16, FP16, FP8
Chassis Form FactorHGX-style OAM server
Software EcosystemIntel Gaudi SynapseAI SDK, PyTorch, TensorFlow, Optimum Habana
Framework SupportPyTorch, TensorFlow, Hugging Face Transformers (via Optimum Habana)
Cooling SupportAir cooling / Direct Liquid Cooling (DLC) depending on configuration
Target WorkloadsAI training, LLM fine-tuning, generative AI inference, HPC

Frequently Asked Questions about Quanta Computer Intel Gaudi 3 8-OAM HGX-Style Server QCT-G3

What server platforms accept the Quanta Computer Intel Gaudi 3 8-OAM HGX-Style Server QCT-G3?

Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.

How long is the lead time on AI GPUs?

Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.

Do you supply matched networking (Quantum InfiniBand / Spectrum-X)?

Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.

Can you help with NVIDIA AI Enterprise licensing?

Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.