AMD released model weight profiles on September 25, 2026, saying it provides detailed breakdowns of how parameters are allocated across different model components. It is the company's first AI/ML update since the launch of Llama-3.1.

AMD reported 8 billion parameters in the Llama-3.1-8B-Instruct-FP8-KV model, measured on the FP8 quantization format. That compares with 9.08 GB of disk space used by the model checkpoint, which includes BF16 parameters.

The Llama-3.1 model is built on AMD's Quark quantization technology and targets efficient inference for large language models. Availability begins with the model's release on September 25, 2026, initially for developers and researchers.

"Model weight profiles show how many model parameters are devoted to embeddings, attention, dense layers, and other components," said Dominic Widdows, AMD's blog author. These profiles help explain discrepancies between parameter counts and disk space usage.

The announcement follows AMD's earlier work on quantization techniques for AI models. AMD said these profiles are important as lower-bit formats become more common, enabling better compression and performance.

AMD did not say how these profiles will be used beyond model optimization, and it raised the question of how different quantization recipes affect model performance. The company said it will continue to release more details on model architecture.

Source: amd