Brand: AMD | Category: GPUs
SKU: AS-8125GS-TNMR2 | Part #: AS-8125GS-TNMR2 | MPN: AS-8125GS-TNMR2
Contact for Pricing — Request a Quote
The Supermicro AS-8125GS-TNMR2 is an 8U rackmount server platform engineered around eight AMD Instinct MI300X OAM (Open Accelerator Module) GPUs, delivering a unified compute-and-memory architecture purpose-built for large-scale AI training, inference, and HPC workloads. Each MI300X accelerator integrates 192 GB of HBM3 memory with 5.3 TB/s of aggregate memory bandwidth per GPU, and the eight-GPU configuration provides a total of 1,536 GB of pooled GPU memory within a single server node—enabling models and datasets that would otherwise require multi-node distribution to reside entirely in-node.
The MI300X is built on AMD's CDNA 3 GPU architecture, combining CPU-class compute dies with stacked HBM3 memory in a single package. Each GPU delivers up to 1,307 TFLOPS of FP8 peak throughput and 653 TFLOPS of FP16, making the platform well-suited for generative AI inference at scale, large language model (LLM) fine-tuning, and dense matrix workloads. The Supermicro AS-8125GS-TNMR2 server chassis provides the thermal, power, and interconnect infrastructure required to sustain eight MI300X OAMs under full load, including high-bandwidth NVLink-class peer-to-peer GPU communication through AMD's Infinity Fabric interconnect.
Designed for enterprise datacenters and hyperscale AI infrastructure, the AS-8125GS-TNMR2 supports AMD ROCm open software platform compatibility, enabling integration with popular AI frameworks including PyTorch, TensorFlow, and JAX. The platform is relevant to organizations in financial services, healthcare AI, national research computing, telecommunications, and cloud service providers deploying inference endpoints or training clusters at scale. Omnixon Global supplies this platform to enterprise and institutional buyers across the UAE, GCC, EMEA, and APAC regions.
| Manufacturer | Supermicro |
| GPU Brand | AMD |
| GPU Model | AMD Instinct MI300X OAM |
| GPU Architecture | AMD CDNA 3 |
| Number of GPUs | 8 |
| GPU Memory per Accelerator | 192 GB HBM3 |
| Total GPU Memory (8x) | 1,536 GB HBM3 |
| Memory Bandwidth per GPU | 5.3 TB/s |
| Peak FP8 Throughput (per GPU) | 1,307 TFLOPS |
| Peak FP16 Throughput (per GPU) | 653 TFLOPS |
| Peak FP64 Throughput (per GPU) | 163.4 TFLOPS |
| GPU Interconnect | AMD Infinity Fabric (GPU-to-GPU) |
| Form Factor | 8U Rackmount |
| Manufacturer Part Number | AS-8125GS-TNMR2 |
| GPU Form Factor | OAM (Open Accelerator Module) |
| Software Platform | AMD ROCm (open-source) |
| Target Workloads | Generative AI, LLM training and inference, HPC, scientific computing |
| Supported Frameworks | PyTorch, TensorFlow, JAX (via ROCm) |
Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.
| Brand | AMD |
| Category | GPUs |
| SKU | AS-8125GS-TNMR2 |
| Part Number | AS-8125GS-TNMR2 |
| Condition | New |
| GPU Brand | AMD |
| GPU Model | AMD Instinct MI300X OAM |
| GPU Architecture | AMD CDNA 3 |
| Number of GPUs | 8 |
| GPU Memory per Accelerator | 192 GB HBM3 |
| Total GPU Memory (8x) | 1,536 GB HBM3 |
| Memory Bandwidth per GPU | 5.3 TB/s |
| Peak FP8 Throughput (per GPU) | 1,307 TFLOPS |
| Peak FP16 Throughput (per GPU) | 653 TFLOPS |
| Peak FP64 Throughput (per GPU) | 163.4 TFLOPS |
| GPU Interconnect | AMD Infinity Fabric (GPU-to-GPU) |
| Form Factor | 8U Rackmount |
| Manufacturer Part Number | AS-8125GS-TNMR2 |
| GPU Form Factor | OAM (Open Accelerator Module) |
| Software Platform | AMD ROCm (open-source) |
| Target Workloads | Generative AI, LLM training and inference, HPC, scientific computing |
| Supported Frameworks | PyTorch, TensorFlow, JAX (via ROCm) |
Reference servers include Dell PowerEdge XE9680 / XE9712, HPE Cray XD670, Lenovo ThinkSystem SR685a / SR675 V3, Supermicro AS-A21GE / SYS-821GE, Gigabyte G593 / G894, ASUS ESC. Share your target platform in the RFQ and we will confirm chassis-to-GPU compatibility and recommended NIC pairing.
Highly model-dependent. L40S / RTX-class: typically 3-6 weeks. H100/H200/B200 in SXM form factor: 12-16 weeks for whole-platform allocations. We quote genuine-channel ETAs only — no grey-market promises.
Yes — Omnixon stocks the full NVIDIA networking lineup (Quantum-2 / Quantum-X InfiniBand, Spectrum-X Ethernet, ConnectX NICs, BlueField DPUs) so we can quote a complete training-cluster BOM, not just the GPUs.
Yes. We hold genuine-channels for NVIDIA AI Enterprise software subscriptions. Add it to your RFQ and we quote node-aligned licensing along with the hardware.