AMD has demonstrated efficient inference capabilities for the MiniMax-M3 model on its Instinct MI355X GPUs using the ATOM inference engine and ATOMesh for multi-node orchestration. The model, released in June 2026, features a 1-million-token context window and native support for text, image, and video understanding. The performance benchmarks highlight AMD's ability to serve large-scale models efficiently, with results from the SemiAnalysis InferenceX open-source benchmark platform.

The single-node performance tests showed that the MI355X with ATOM and EAGLE3 speculative decoding achieved 340–370 tok/s/user at high interactivity, outperforming NVIDIA's B200 with vLLM in certain metrics. Additionally, the MI355X delivered ~0.6k tok/s/GPU throughput in FP4, compared to ~0.2k tok/s/GPU for the B200. In multi-node configurations, the MI355X with ATOMesh and Mooncake for KV cache transfer achieved 0.3k–5k tok/s/GPU throughput, surpassing NVIDIA’s B300 with Dynamo vLLM.

AMD's optimization efforts, including EAGLE3 speculative decoding and AITER-optimized kernels, have played a critical role in enhancing the performance of MiniMax-M3 on Instinct GPUs. The company's early lead in the optimization race has been maintained through continuous iterations and improvements to the inference stack. The results underscore the competitive landscape in AI inference and the ongoing efforts to improve serving efficiency for large models.

Source: amd