Supermicro A+ Server AS-8125GS-TNMR2 8U GPU Server with AMD Instinct MI300X 8-GPU

Supermicro A+ Server AS-8125GS-TNMR2 8U GPU Server with AMD Instinct MI300X 8-GPU

Brand: AMD | Category: GPUs

SKU: AS-8125GS-TNMR2 | Part #: AS-8125GS-TNMR2 | MPN: AS-8125GS-TNMR2

Contact for Pricing — Request a Quote

Request a Quote Contact Us

About the Supermicro A+ Server AS-8125GS-TNMR2 8U GPU Server with AMD Instinct MI300X 8-GPU

The Supermicro AS-8125GS-TNMR2 is an 8U rackmount server platform engineered around eight AMD Instinct MI300X OAM (Open Accelerator Module) GPUs, delivering a unified compute-and-memory architecture purpose-built for large-scale AI training, inference, and HPC workloads. Each MI300X accelerator integrates 192 GB of HBM3 memory with 5.3 TB/s of aggregate memory bandwidth per GPU, and the eight-GPU configuration provides a total of 1,536 GB of pooled GPU memory within a single server node—enabling models and datasets that would otherwise require multi-node distribution to reside entirely in-node.

The MI300X is built on AMD's CDNA 3 GPU architecture, combining CPU-class compute dies with stacked HBM3 memory in a single package. Each GPU delivers up to 1,307 TFLOPS of FP8 peak throughput and 653 TFLOPS of FP16, making the platform well-suited for generative AI inference at scale, large language model (LLM) fine-tuning, and dense matrix workloads. The Supermicro AS-8125GS-TNMR2 server chassis provides the thermal, power, and interconnect infrastructure required to sustain eight MI300X OAMs under full load, including high-bandwidth NVLink-class peer-to-peer GPU communication through AMD's Infinity Fabric interconnect.

Designed for enterprise datacenters and hyperscale AI infrastructure, the AS-8125GS-TNMR2 supports AMD ROCm open software platform compatibility, enabling integration with popular AI frameworks including PyTorch, TensorFlow, and JAX. The platform is relevant to organizations in financial services, healthcare AI, national research computing, telecommunications, and cloud service providers deploying inference endpoints or training clusters at scale. Omnixon Global supplies this platform to enterprise and institutional buyers across the UAE, GCC, EMEA, and APAC regions.

Ideal for

  • Large language model (LLM) inference serving, where the 1,536 GB total GPU memory pool enables full in-node hosting of 70B–400B parameter models without tensor sharding across nodes
  • Generative AI model fine-tuning and instruction tuning for enterprise-customized foundation models in regulated industries such as finance and healthcare
  • High-performance scientific computing and simulation workloads including molecular dynamics, climate modeling, and seismic processing that require sustained FP64 throughput
  • AI-driven drug discovery and genomics pipelines requiring large memory capacity for simultaneous processing of high-dimensional biological datasets
  • Enterprise AI inference consolidation, replacing multi-node GPU clusters with a single 8U server node to reduce rack footprint, interconnect complexity, and operational overhead
  • Sovereign AI and national research computing deployments where a single dense node must support multiple concurrent research teams running independent large-model workloads

Technical specifications

ManufacturerSupermicro
GPU BrandAMD
GPU ModelAMD Instinct MI300X OAM
GPU ArchitectureAMD CDNA 3
Number of GPUs8
GPU Memory per Accelerator192 GB HBM3
Total GPU Memory (8x)1,536 GB HBM3
Memory Bandwidth per GPU5.3 TB/s
Peak FP8 Throughput (per GPU)1,307 TFLOPS
Peak FP16 Throughput (per GPU)653 TFLOPS
Peak FP64 Throughput (per GPU)163.4 TFLOPS
GPU InterconnectAMD Infinity Fabric (GPU-to-GPU)
Form Factor8U Rackmount
Manufacturer Part NumberAS-8125GS-TNMR2
GPU Form FactorOAM (Open Accelerator Module)
Software PlatformAMD ROCm (open-source)
Target WorkloadsGenerative AI, LLM training and inference, HPC, scientific computing
Supported FrameworksPyTorch, TensorFlow, JAX (via ROCm)

Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.

Technical Specifications

BrandAMD
CategoryGPUs
SKUAS-8125GS-TNMR2
Part NumberAS-8125GS-TNMR2
ConditionNew
GPU BrandAMD
GPU ModelAMD Instinct MI300X OAM
GPU ArchitectureAMD CDNA 3
Number of GPUs8
GPU Memory per Accelerator192 GB HBM3
Total GPU Memory (8x)1,536 GB HBM3
Memory Bandwidth per GPU5.3 TB/s
Peak FP8 Throughput (per GPU)1,307 TFLOPS
Peak FP16 Throughput (per GPU)653 TFLOPS
Peak FP64 Throughput (per GPU)163.4 TFLOPS
GPU InterconnectAMD Infinity Fabric (GPU-to-GPU)
Form Factor8U Rackmount
Manufacturer Part NumberAS-8125GS-TNMR2
GPU Form FactorOAM (Open Accelerator Module)
Software PlatformAMD ROCm (open-source)
Target WorkloadsGenerative AI, LLM training and inference, HPC, scientific computing
Supported FrameworksPyTorch, TensorFlow, JAX (via ROCm)

Frequently Asked Questions about Supermicro A+ Server AS-8125GS-TNMR2 8U GPU Server with AMD Instinct MI300X 8-GPU

What server platforms accept the AMD Instinct MI300X 8-GPU OAM Server Node (Supermicro AS-8125GS-TNMR2)?

Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.

How long is the lead time on AI GPUs?

Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.

Do you supply matched networking (Quantum InfiniBand / Spectrum-X)?

Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.

Can you help with NVIDIA AI Enterprise licensing?

Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.