AMD released Zebra-HyLo on September 22, 2026, saying it enables long-context hybrid language models without requiring full pretraining. It is the company's first research update in the artificial intelligence field since the release of the Instinct GPU line.

AMD reported models using Zebra-HyLo achieved up to 32x context extension, measured on long-context benchmarks like RULER. That compares with previous methods that scored 55.1 at 8K and 0.8 at 64K.

Zebra-HyLo is built on AMD Instinct™ MI300X GPUs and targets efficient long-context language processing. Availability begins with open-source checkpoints, initially for researchers and developers.

"Zebra-HyLo treats long-context preservation as a first-class training objective rather than something the architecture provides for free," said Parsa Ashrafi Fashi, lead researcher. The method delivers models that extend usable context by up to 32x, cut KV-cache memory by more than 90%, and serve up to 2M tokens in vLLM on eight AMD Instinct™ MI300X GPUs.

The announcement follows AMD's focus on hybrid architectures for AI. AMD said the approach is significant for enabling long-context models without requiring full pretraining, which aligns with the company's strategy for efficient AI development.

AMD did not say how the method will scale to even longer contexts, and raised the open question of whether further optimization is needed for real-world applications. The company said it will continue to refine the method for broader use.

Source: amd