AMD released a disaggregated inference solution for the Kimi-K3 MXFP4 model on October 8, 2026, saying it enables efficient serving of a 2.8-trillion-parameter model across multiple GPUs. It is the company's first hardware update for AI inference since the release of the MI300 series in 2025.
AMD reported a memory footprint of 1453.7 GiB for the Kimi-K3 MXFP4 model, measured on an eight-GPU MI300X node. That compares with a usable high-bandwidth memory capacity of 1504 GiB, leaving 50 GiB for key-value (KV) cache, activations, and communication heap.
The Kimi-K3 MXFP4 model is built on the CDNA 3 (gfx942) architecture and targets large-scale language model serving. Availability begins with the MI300X and MI325X GPUs, initially for enterprise and cloud service providers.
"MXFP4, the quantization format that makes a 2.8-trillion-parameter Mixture-of-Experts (MoE) model tractable, has no native matrix-multiply instruction on the CDNA 3 (gfx942) architecture," said Ravi Gupta, AMD's senior director of AI/ML. The company's solution requires requantizing experts into a format the matrix engine accelerates.
The announcement follows the release of the AMD Disaggregated Inference (DI) Blog Series, which takes open frontier models from single-node demos to multi-node serving on AMD Instinct GPUs. AMD framed the significance as a step toward scalable, production-ready AI inference solutions.
AMD did not say how the MXFP4 format will perform on future architectures, and the company raised the question of whether the memory savings from quantization will survive the requantization process. The company said it will continue to optimize the deployment shape for different models.
Source: amd