2.4B Open Vision Model Released
Key point
CohereLabs released a 2.4B vision-language model under the Apache 2.0 license.
Details
CohereLabs released North Micro Vision Instruct, an open-weight vision-language model with 2.4B parameters.
The model supports native-resolution processing to preserve image aspect ratios and details, targeting the following tasks:
- Visual Question Answering (VQA) and image captioning
- Visual grounding and spatial understanding
- OCR, chart and document understanding, and structured information extraction
- Multilingual and multi-image understanding
The full architecture consists of a 2B parameter language model and a 400M parameter vision encoder. It processes mixed text and image inputs and supports multiple languages, including English, Korean, Chinese, and Japanese.
The language model has a context window of 128K tokens, but the validated range for multimodal inputs is up to 8K tokens. The weights are provided under the Apache 2.0 license, allowing use for prototyping, task-specific fine-tuning, and developing lightweight multimodal applications.
However, it is a base model for customization rather than a replacement for large general-purpose chatbots, and as it is not a dedicated reasoning model, it has limitations in mathematics and code generation capabilities.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.