IBM has launched the latest iteration of its Granite large language models, designed for local deployment and self-hosting. The Granite 4.2 series includes variants with 3B, 8B, and 30B parameters, each offering a 128,000-token context window. These models are built using a decoder-only architecture, similar to previous versions in the series. IBM emphasized that this release focuses on reasoning capabilities, a key differentiator in the local model space.
The 8B and 30B variants include an agentic reinforcement-learning block, enabling expanded functions like terminal use, web searching, and external tool integration. The 3B model also supports tools but lacks the same level of specialized training. IBM stated that the reasoning focus means models can handle complex tasks through chain-of-thought reasoning, though this may result in slower response times and higher compute demands. The company noted that this approach prioritizes rigorous and accurate responses over speed in certain scenarios.
IBM’s Granite models have traditionally focused on enterprise use cases rather than speed or innovation. This release aligns with growing interest in local models as a cost-effective alternative to cloud-based systems. The trend is driven by both developers and enterprises seeking to reduce costs and dependency on cloud services. Model routers, which direct tasks to appropriate models, are also gaining popularity for their efficiency and flexibility. IBM’s strategy reflects a broader industry shift toward localized AI solutions.
Source: arstechnica