AMD released an update to its GLM-5.2-MXFP4 model on AMD Instinct™ MI355X GPUs, enabling prefill context parallelism. It is the company's first software update in the AI/ML category since its initial release.

AMD reported a 43% to 54% increase in total throughput on 40,960- and 61,440-token prompts at concurrency 16 and above, measured on SGLang with --tp-size 4. That compares with an 8% to 18% decrease in throughput on 1024-token prompts.

The model is built on the SGLang framework and targets large-scale language processing tasks. Availability begins with the current SGLang build, initially for developers and researchers. "Prefill context parallelism assigns prompt tokens round robin across four ranks," said Bobo Fang, lead researcher.

This approach allows each rank to process one quarter of the sequence. The announcement follows AMD's ongoing efforts to optimize AI workloads on its GPUs.

AMD did not say how the model performs on shorter prompts, and raised the question of when the split is wrong. The company noted that the practical fallback will be covered near the end of the post.

Source: amd