Cohere announced the release of North-Micro-Vision-Instruct, a 2.4B-parameter open-weight vision-language model (VLM) with native-resolution image support. The model is designed for compact, specialized multimodal applications and is the smallest VLM developed by the company to date. It combines broad image-understanding capabilities with native-resolution support, making it suitable for fine-tuning across various visual domains and workflows. With appropriate inference stacks and quantization, the model can also run on laptops and edge or mobile-class hardware. The release highlights Cohere's commitment to sovereign AI through open weights, clear licensing, and transparent evaluation.
North-Micro-Vision-Instruct is built to balance broad visual capabilities with a size suitable for local, edge-aware, and specialized deployments. It preserves the aspect ratio and fine detail of documents, tables, charts, screenshots, and forms instead of reducing inputs to a small square. The model supports multilingual visual understanding, covering multiple languages and visual domains, including documents, charts, and natural images. It also provides a compact foundation for adaptation to domain-specific data, tasks, and deployment constraints. The architecture includes a native-resolution vision encoder, a projector, and a compact language model, all working together to enhance visual and linguistic comprehension.
The model's development reflects Cohere's broader efforts in sovereign AI, emphasizing model development alongside licensing, open weights, and transparent evaluation. The release includes public model weights and vLLM support, which is expected soon. The training process involved four stages, with Stage 2 split into two resolution phases, focusing on increasing resolution while jointly training the encoder, projector, and language model. The final stage incorporated a simplified variant of Mixed Preference Optimization to improve safety, formatting, and response quality.
Source: huggingface