Cohere Unveils Multilingual VLM 'Aya Vision'
·2025.03.04 09:00
Key point
Cohere For AI has released the 'Aya Vision' series, an open-weight multimodal model supporting 23 languages.
1 / 2
Details
Cohere For AI has unveiled the Aya Vision model family (8B, 32B), which possesses strong language and vision understanding capabilities across 23 languages.
Key Performance and Features:
- Aya Vision 32B: Demonstrates performance that surpasses models more than twice its size, such as Llama-3.2 90B Vision, Molmo 72B, and Qwen2.5-VL 72B.
- Aya Vision 8B: Proved its efficiency by recording overwhelming win rates against similarly-sized parameter models such as Qwen2.5-VL 7B and Gemini Flash 1.5 8B.
- Multilingual Support: Delivers excellent performance in image captioning, visual question answering (VQA), text generation, and translation tasks across 23 languages.
Technical Innovations:
- Uses SigLIP2 as the vision encoder initialization model.
- Applied Dynamic Resizing & Tiling technology for processing high-resolution images.
- Improved latency and throughput by compressing the number of image tokens by 4x through the Pixel Shuffle downsampling technique.
Sharing Research Resources: A new benchmark, AyaVisionBench, and a multilingual version of the mWildVision dataset have been released together to support multimodal research.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.