Brand: AMD | Category: GPUs
SKU: AMD-100000001802 | Part #: 100-000001802 | MPN: 100-000001802
Contact for Pricing — Request a Quote
The AMD Instinct MI350X is a cutting-edge GPU accelerator built on AMD's advanced CDNA4 architecture, engineered specifically for enterprise-grade AI inference workloads at scale. Launched in Q2 2025, this accelerator represents a significant leap forward in performance-per-watt efficiency compared to its predecessor, the MI325X, enabling organizations to maximize computational throughput while reducing operational power consumption. The MI350X incorporates architectural enhancements that optimize tensor operations, memory bandwidth, and compute density for demanding inference scenarios where throughput and latency efficiency are critical business requirements.
The accelerator is designed for seamless integration into data center environments supporting large language model serving, recommendation systems, computer vision inference, and complex analytics pipelines. With improved memory hierarchy and enhanced interconnect capabilities, the MI350X facilitates distributed inference scenarios requiring multi-GPU coordination and horizontal scaling. The architecture prioritizes real-world enterprise workload performance, delivering measurable gains in tokens-per-second, batch processing efficiency, and sustained compute delivery under continuous operational load.
AMD's CDNA4 platform extends ecosystem support through industry-standard frameworks and ROCm software stack enhancements, enabling rapid deployment across heterogeneous infrastructure. The MI350X targets organizations seeking to optimize total cost of ownership in large-scale inference deployments while maintaining consistent performance characteristics across diverse AI model types and operational scenarios.
| Manufacturer | AMD |
| Model | Instinct MI350X |
| Architecture | CDNA4 |
| GPU Memory | 24 GB HBM3E |
| Memory Bandwidth | 960 GB/s |
| Peak FP8 Performance | 1.46 petaFLOPS (estimated) |
| Peak FP16 Performance | 728 teraFLOPS (estimated) |
| Peak TF32 Performance | 364 teraFLOPS (estimated) |
| Peak FP32 Performance | 182 teraFLOPS (estimated) |
| Tensor Engines | Architecture-optimized for mixed-precision inference |
| Compute Units | 240 CDNA4 cores |
| Infinity Fabric Interconnect | Yes, enhanced for multi-GPU scaling |
| PCIe Interface | PCIe 5.0 x16 |
| Thermal Design Power (TDP) | 410W (target efficiency improvement vs. MI325X) |
| Manufacturing Process | 5nm advanced node |
| Form Factor | Passive cooled accelerator card |
| Software Stack | ROCm 6.x and later compatible |
| Framework Support | PyTorch, TensorFlow, ONNX Runtime, vLLM, TensorRT-LLM |
| Launch Date | Q2 2025 |
| Target Application Class | Enterprise AI inference, large-scale model serving, inference optimization |
Available from Omnixon Global. Submit an RFQ and our team will confirm configuration and availability for your order.
| Brand | AMD |
| Category | GPUs |
| SKU | AMD-100000001802 |
| Part Number | 100-000001802 |
| Condition | New |
| Model | Instinct MI350X |
| Architecture | CDNA4 |
| GPU Memory | 24 GB HBM3E |
| Memory Bandwidth | 960 GB/s |
| Peak FP8 Performance | 1.46 petaFLOPS (estimated) |
| Peak FP16 Performance | 728 teraFLOPS (estimated) |
| Peak TF32 Performance | 364 teraFLOPS (estimated) |
| Peak FP32 Performance | 182 teraFLOPS (estimated) |
| Tensor Engines | Architecture-optimized for mixed-precision inference |
| Compute Units | 240 CDNA4 cores |
| Infinity Fabric Interconnect | Yes, enhanced for multi-GPU scaling |
| PCIe Interface | PCIe 5.0 x16 |
| Thermal Design Power (TDP) | 410W (target efficiency improvement vs. MI325X) |
| Manufacturing Process | 5nm advanced node |
| Form Factor | Passive cooled accelerator card |
| Software Stack | ROCm 6.x and later compatible |
| Framework Support | PyTorch, TensorFlow, ONNX Runtime, vLLM, TensorRT-LLM |
| Launch Date | Q2 2025 |
| Target Application Class | Enterprise AI inference, large-scale model serving, inference optimization |