AMD released Quark on September 25, 2026, saying it provides memory-efficient quantization workflows for local AI deployment. It is the company's first software update in the AI serving category since the Ryzen AI Max processor launch.
AMD reported a 30% reduction in model size, measured on the Strix Halo system, reducing the model from roughly 70 GB in 16-bit precision to about 21 GB. That compares with the original size of the 35B-parameter Qwen3.6-35B-A3B model.
Quark is built on the AMD ROCm platform and targets local AI deployment on devices like Strix Halo. Availability begins with the release of the Quark software, initially for developers and researchers.
"Applications & models LLM, Optimization, Serving AI, Data Science, Systems For local AI, running a model on the target device is only part of the deployment workflow," said HongWei Meng, lead author of the post. Model preparation, including quantization and export, is another important step.
The announcement follows AMD's Ryzen AI Max processor launch. AMD framed the significance of Quark as an extension of its AI inference work, covering quantization, export, and application-level deployment on Strix Halo.
AMD did not say how the Quark quantization will perform on other platforms, and the open question is whether the same efficiency can be achieved on different hardware. The source says the quantized models used in the post are available on Hugging Face.
Source: amd