NVIDIA GH200 Grace Hopper Superchip 480GB HBM3e

NVIDIA GH200 Grace Hopper Superchip 480GB HBM3e

Brand: NVIDIA | Category: GPUs

SKU: 920-23687-2560-000 | Part #: 920-23687-2560-000 | MPN: 920-23687-2560-000

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the NVIDIA GH200 Grace Hopper Superchip 480GB HBM3e

The NVIDIA GH200 Grace Hopper Superchip combines the NVIDIA Grace CPU and Hopper GPU architecture on a single module, interconnected via NVLink-C2C with 900 GB/s of chip-to-chip bandwidth. This unified memory architecture enables the CPU and GPU to share a coherent memory space, dramatically reducing data movement overhead for memory-intensive AI and HPC workloads. The 480 GB HBM3e configuration delivers up to 4.9 TB/s of GPU memory bandwidth, making it one of the highest-bandwidth compute platforms available for enterprise datacenter deployment.

Built on the Hopper GPU architecture, the GH200 incorporates fourth-generation Tensor Cores, second-generation Transformer Engine, and NVLink 4.0 fabric support for multi-GPU scaling. The Grace CPU component features 72 Arm Neoverse V2 cores and 480 GB of LPDDR5X system memory, providing a tightly integrated compute substrate well-suited to large-scale transformer model inference and training. The combination of HBM3e GPU memory and LPDDR5X CPU memory gives the full superchip 624 GB of total addressable coherent memory.

The GH200 480 GB HBM3e targets enterprises running generative AI model training, large language model inference, scientific simulation, and high-performance computing workloads at scale. It is designed for deployment in high-density datacenter nodes such as the NVIDIA MGX reference architecture and DGX GH200 systems, and supports NVLink Switch System configurations for multi-node GPU clusters. The module is qualified for demanding enterprise AI infrastructure requiring maximum memory capacity, bandwidth, and compute density in a single accelerator package.

Ideal for

  • Large language model training and fine-tuning for enterprise AI platforms requiring models exceeding 70 billion parameters
  • Generative AI inference serving for production deployments where large model weights must reside entirely in accelerator memory
  • High-performance computing simulations in computational fluid dynamics, molecular dynamics, and climate modeling demanding extreme memory bandwidth
  • Recommender system training on massive embedding tables that benefit from the unified 624 GB coherent CPU and GPU memory space
  • Multi-modal AI workload development combining vision, language, and structured data models within a single high-bandwidth memory domain
  • Sovereign AI datacenter buildouts across UAE, GCC, EMEA, and APAC requiring dense, scalable GPU compute infrastructure

Technical specifications

ManufacturerNVIDIA
Manufacturer Part Number920-23687-2560-000
Product FamilyGrace Hopper Superchip
GPU ArchitectureNVIDIA Hopper (GH200)
CPU ArchitectureNVIDIA Grace (72x Arm Neoverse V2 cores)
GPU Memory480 GB HBM3e
GPU Memory Bandwidth4.9 TB/s
CPU Memory480 GB LPDDR5X
Total Coherent Memory624 GB
CPU-GPU InterconnectNVLink-C2C, 900 GB/s bidirectional
FP8 Tensor Core PerformanceUp to 8 PFLOPS (Transformer Engine with sparsity)
FP16 Tensor Core PerformanceUp to 4 PFLOPS
TF32 Tensor Core PerformanceUp to 2 PFLOPS
NVLink VersionNVLink 4.0
PCIe InterfacePCIe Gen5 x16
Tensor CoresFourth-generation
Transformer EngineSecond-generation
Form FactorSXM5 Module (NVIDIA MGX compatible)
Thermal Design Power (TDP)900 W

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandNVIDIA
CategoryGPUs
SKU920-23687-2560-000
Part Number920-23687-2560-000
ConditionNew
Manufacturer Part Number920-23687-2560-000
Product FamilyGrace Hopper Superchip
GPU ArchitectureNVIDIA Hopper (GH200)
CPU ArchitectureNVIDIA Grace (72x Arm Neoverse V2 cores)
GPU Memory480 GB HBM3e
GPU Memory Bandwidth4.9 TB/s
CPU Memory480 GB LPDDR5X
Total Coherent Memory624 GB
CPU-GPU InterconnectNVLink-C2C, 900 GB/s bidirectional
FP8 Tensor Core PerformanceUp to 8 PFLOPS (Transformer Engine with sparsity)
FP16 Tensor Core PerformanceUp to 4 PFLOPS
TF32 Tensor Core PerformanceUp to 2 PFLOPS
NVLink VersionNVLink 4.0
PCIe InterfacePCIe Gen5 x16
Tensor CoresFourth-generation
Transformer EngineSecond-generation
Form FactorSXM5 Module (NVIDIA MGX compatible)
Thermal Design Power (TDP)900 W

Frequently Asked Questions about NVIDIA GH200 Grace Hopper Superchip 480GB HBM3e

What server platforms accept the NVIDIA GH200 Grace Hopper Superchip 480GB HBM3e?

Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.

How long is the lead time on AI GPUs?

Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.

Do you supply matched networking (Quantum InfiniBand / Spectrum-X)?

Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.

Can you help with NVIDIA AI Enterprise licensing?

Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.