AMD released Kimi-K3 benchmark results on MI350X GPUs, saying it tested the model's performance across three serving frameworks. It is the company's first software update since the release of Kimi-K2.
AMD reported 0.5% performance variance across vLLM, SGLang, and ATOM on an 8× MI350X node, measured with 8192 input and 1024 output tokens. That compares with a wider spread at lower and higher concurrency levels.
Kimi-K3 is built on MXFP4 Mixture-of-Experts (MoE) architecture and targets large-scale AI inference workloads. Availability begins with day-0 support for AMD Instinct™ GPUs, initially for developers and enterprise users.
"The answer used here is MAD (Model Automation and Dashboarding), AMD’s open-source benchmarking harness for AMD Instinct™ GPUs," said Yu Shao. MAD keeps a declarative model registry, models.json, in which each model is described with its parameters and dependencies.
The announcement follows Moonshot AI’s release of Kimi-K3 weights on July 27, 2026. AMD emphasized the importance of the model’s hybrid attention mechanism and native vision capabilities.
AMD did not say how the model will perform on multi-node setups, and raised the open question of how different frameworks will handle the model’s unique attention mechanisms. The source says the benchmark results are a starting point for further optimization.
Source: amd