Microsoft has released MAI Code 1.1 Flash, a code model for GitHub Copilot. The model claims to write better code, be 25 percent more token-efficient, and cost a quarter of its June predecessor. Developers accepted 4 percent more of its output. Training involved 'hundreds of thousands of reinforcement-learning environments in GitHub Copilot.'
In benchmarks, MAI-Code-1.1-Flash edges past its predecessor and mini-models from Anthropic and OpenAI but gets crushed by Deepseek-V4-Flash-0731. On the SWE-bench, it scored 72.6%, compared to 71.6% for MAI-Code-1-Flash, 69.8% for Haiku 4.5, 69.2% for GPT-5.4 mini, and no published result for Deepseek-V4-Flash-0731. On the Terminal Bench, it scored 2.1, compared to 62.9% for MAI-Code-1-Flash, 51.7% for Haiku 4.5, 49.4% for GPT-5.4 mini, and 82.7% for Deepseek-V4-Flash-0731. Cost per token doesn't tell the whole story without factoring in usage efficiency, but the gap in Deepseek's favor is likely significant either way.
Microsoft's open AI talk doesn't match its model strategy. None of this squares with Microsoft's recent push to paint itself as an open AI champion. Instead of tapping more capable, freely available alternatives like Deepseek-V4-Flash, the company is sinking resources into a weaker, pricier in-house model that's proprietary and likely won't get an open-weights release. The reason is probably the same one behind Microsoft's recent Copilot shakeup, where it swapped out OpenAI and Anthropic models for its own cheaper MAI alternatives to cut costs. The trade-off was worse performance for better margins.
Source: thedecoder