Alibaba's Qwen3.8 Max has achieved a score of 56 on the Artificial Analysis Intelligence Index, according to Artificial Analysis. This marks a 10-point increase from Qwen3.7 Max's score of 46. The model now matches the performance of Claude Opus 4.8 and outperforms GLM-5.2, which scored 51. However, it still lags behind Kimi K3, which achieved a score of 57. Notably, Kimi K3 also operates at 25 percent lower cost than Qwen3.8 Max.

The model's performance improvement comes at a cost. Qwen3.8 Max requires 64 steps per task instead of 14, and input tokens have increased 15 times due to the model resending the full conversation history at each step. This results in slower processing and higher costs. Despite lower token prices, the price-to-performance ratio for Alibaba's model has worsened. A single task on the Intelligence Index now costs $1.14, more than double the $0.53 for Qwen3.7 Max. Kimi K3, meanwhile, costs $0.86 per task and scores one point higher.

According to the source, there are also regressions compared to the previous version. The AA-LCR test score dropped by 2 points, indicating reduced ability to correctly pull information from long texts. The AA-Omniscience score fell by 10 points, reflecting a higher hallucination rate from 23 to 40 percent. The accuracy rate remains around 31 percent, but the model now guesses more frequently instead of admitting it doesn't know. Source: thedecoder