AMD released Primus, a framework for training hybrid large language models on Instinct GPUs, saying it enables pre-training hybrid models from scratch. It is the company's first research update in AI/ML since the launch of Instinct MI355X GPUs.
AMD reported final loss of 3.382 on a 300M parameter model, measured on sequence length 2048 in bfloat16. That compares with 3.411 on a Pure GDN 300M model.
Primus is built on AMD Instinct MI355X GPUs and targets hybrid model training with linear-recurrent mixers. Availability begins with open-source configurations, initially for researchers and developers.
"We trained five models to completion on one node of eight MI355X GPUs," said Vansh Bhatia, lead author. The models included Pure GDN, Pure KDA, and 75% hybrid variants.
The announcement follows the release of FineWeb-Edu, a dataset used for training. AMD said the framework supports hybrid models by combining attention layers with linear-recurrent blocks.
AMD did not say how the framework will scale to larger models, and raised the open question of how different mixers affect training quality. The source says the configurations used are included in the repository.
Source: amd