Z.ai has released GLM-5.3-Flash, a model with 320 billion parameters and a context window of one million tokens. According to Artificial Analysis, the model scores 57 on the Intelligence Index at maximum reasoning effort, just three points behind the larger GLM-5.3, which scores 60. The model also matches GPT-5.6 Terra and Muse Spark 1.2 in performance. It runs entirely on Chinese AI chips and is priced at 0.09 dollars per task, making it significantly more affordable than its predecessor, GLM-5.3, which costs 0.68 dollars per task.

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, according to Z.ai. It ships under an MIT license and has its weights available on Hugging Face. On Z.ai's API, the model costs 0.15 dollars per million input tokens and 0.50 dollars per million output tokens, a little over ten percent of the price of GLM-5.3. It performs well on agentic tasks, achieving an Elo score of about 1770 on GDPval-AA v2, matching GLM-5.3 and Grok 4.6, and trailing only Claude Opus 5. However, it is less token-efficient, with roughly 90 percent of the output tokens used for reasoning.

Before launch, Z.ai tested the model anonymously as 'ox-alpha' on OpenCode and OpenRouter, where it became the most popular model of the week. All traffic ran on Chinese AI chips, according to Z.ai. SemiAnalysis reports that the model served 100 trillion tokens a day, a level of capacity previously thought possible only for frontier labs. Z.ai claims its hardware efficiency and cost per token are on par with common Nvidia GPUs.

Source: thedecoder