IBM has released the Granite 4.2 language model family, which includes variants with 3 billion, 8 billion, and 30 billion parameters. The models are trained from scratch on approximately 15 trillion tokens and support context windows up to 512,000 tokens. According to IBM, the models can switch between 'thinking' and 'non-thinking' modes to manage computational resources per task. A 'low-effort' mode is designed to conserve resources for simpler queries. The 8B and 30B variants also undergo 'agentic RL' training, allowing them to use tools, write and execute code, and search the web in real sandbox environments. All models support OpenAI-format tool calling and are compatible with vLLM or SGLang. The 30B model performs best in benchmark tests related to agentic tasks and tool use. | Image: IBM

The Granite Speech 5.0 Turbo CTC models, which have 470 million parameters, are twice as fast as previous leaders on the Open ASR Leaderboard, according to IBM. These models can transcribe three hours of audio in one second. All Granite models are available under the Apache 2.0 license on platforms such as Hugging Face, Ollama, GitHub, and others. IBM emphasized that the models are open-weight and accessible for broader use.

Source: thedecoder