Qwen released Qwen3.8-Omni-Flash on September 19, 2026, saying it is the company's first multimodal model built for AI agents. It is the company's first multimodal model since its initial release in 2023.
Qwen reported its model performs on par with Gemini Flash 3.8 in multimodal benchmarks, measured on audio-video tasks. That compares with Gemini Flash 3.8's pricing of $0.75 for input and $3.75 for output per million tokens at its introductory rate.
Qwen3.8-Omni-Flash is built on a multimodal architecture and targets tasks such as video editing, translation, and movie summarization. Availability begins through Qwen Studio, Qwen Cloud, and the API, initially for developers and businesses.
"Qwen3.8-Omni-Flash processes audio and video together, draws conclusions, and uses tools on its own to edit vlogs, translate short videos, or summarize movies," said Matthias Bastian, Qwen's spokesperson. Qwen estimates audio input at under $0.01 per hour, while 720p video with audio at one frame per second runs about $0.20, not counting response costs.
The announcement follows Qwen's earlier release of Qwen-MM-Plugins, which add video editing, speaker recognition, PDF video notes, and reusable workflows to agents like Claude Code, Gemini CLI, and Qwen Code. Qwen did not say when the model will be available for consumer use, and the company raised the question of how the model will scale for more complex tasks.
Source: thedecoder