Google has launched three new Gemini Flash models—3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber—as part of its ongoing efforts to expand its AI offerings. The company emphasized efficiency and cost reduction in these models, with 3.6 Flash showing improvements in token usage and benchmark scores. According to the Artificial Analysis Index, 3.6 Flash is expected to use about 17 percent fewer output tokens than 3.5 Flash, with savings reaching 65 percent on specific benchmarks like DeepSWE. The pricing for 3.6 Flash is set at $1.50 per million input tokens and $7.50 per million output tokens, significantly lower than previous models.

The 3.5 Flash-Lite variant is designed for large workloads with low cost, producing 350 output tokens per second at a rate of $0.30 per million input tokens and $2.50 per million output tokens. It outperforms its predecessor in several coding and agentic benchmarks, including SWE-Bench Pro and OSWorld-Verified. Meanwhile, the 3.5 Flash Cyber model is tailored for cybersecurity tasks and is restricted to governments and partners due to potential risks. It scores 83.2 percent on the CyberGym benchmark, within two points of OpenAI's GPT-5.5-Cyber.

Google's delay in releasing its flagship Gemini 3.5 Pro has raised concerns about its ability to compete with rivals like OpenAI, Anthropic, and Meta, which have already launched more capable frontier models. The company stated that pretraining for Gemini 4 is already underway, calling it its 'most ambitious training run' yet. However, the absence of the Pro model leaves Google without a public model that can compete at the top of the market.

Source: thedecoder