Multimodal Model Visual Salamandra Released
Key point
Visual Salamandra, a 7B-parameter multimodal model, has been released with image and video understanding capabilities.
Details
The Language Technologies Lab has released Visual Salamandra, which can process both images and videos. This model is based on the 7B (7 billion) parameter Salamandra Instructed model, extending multimodal capabilities while maintaining efficiency.
Key Technical Features:
- Uses a Late-fusion architecture combining a SigLIP-So400m encoder with a 2-layer MLP projector.
- Applies a 4-stage training process to align image embeddings with the LLM's latent space.
- Leverages 6.1 million instruction-tuning instances to strengthen Visual Grounding, document understanding, and mathematical reasoning capabilities.
Key Features and Applications:
- VQA (Visual Question Answering): Question answering based on images and videos.
- OCR and Document Understanding: Text extraction and analysis from charts, diagrams, and complex documents.
- Mathematical Reasoning: Solving math problems that include visual information.
- Video Processing: Supports video summarization and event detection capabilities.
In particular, this model focuses on multilingual training that considers the diversity of European languages, contributing to closing the linguistic resource gap found in existing multimodal models.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.