AMD released Primus on September 3, 2026, saying it enables end-to-end DeepSeek-V4-Flash pretraining on AMD Instinct MI355X GPUs. It is the company's first software update in the AI/ML category since the release of DeepSeek-V3 in April 2026.
AMD reported a 10% reduction in single-token inference FLOPs, measured on a one-million-token context, compared to DeepSeek-V3.2. That compares with a 7% reduction in the KV cache.
DeepSeek-V4-Flash is built on the Transformer architecture and targets large-scale language model training. Availability begins with open-source access, initially for developers and researchers.
"DeepSeek-V4-Flash keeps the skeleton you already know from DeepSeek-V3 — a Transformer stack with DeepSeekMoE feed-forward layers and a Multi-Token Prediction head," said Lihuan Zhang, lead researcher at DeepSeek-AI. The model's architecture introduces changes in attention mechanisms and residual connections.
The announcement follows the release of DeepSeek-V4 on April 24, 2026. AMD said the update improves efficiency and scalability for training large models, adding no judgment of its own.
AMD did not say how the model performs on other benchmarks, and it raised the open question of how the sparse attention mechanisms affect training efficiency. The company said the Primus repository includes all configurations and benchmarks for reproducibility.
Source: amd