Ultra-lightweight VLM SmolVLM 256M/500M released
Key point
HuggingFace has released SmolVLM 256M and 500M, ultra-lightweight multimodal models.
Details
HuggingFace announced SmolVLM-256M and SmolVLM-500M, ultra-lightweight Vision Language Models (VLM) that go beyond the existing SmolVLM 2B.
SmolVLM-256M is the world's smallest VLM with 256 million parameters, capable of image captioning, document Q&A, and basic visual reasoning tasks. SmolVLM-500M is a model that balances performance and efficiency, offering higher performance on DocVQA and MMMU, with improved robustness to prompts.
The key technical features are as follows:
- Increased efficiency through adoption of the SigLIP base patch-16/512 vision encoder
- Precise image understanding through processing of larger image resolutions
- Substantial benchmark performance improvements through new tokenization techniques
These models can be loaded immediately via Transformers, MLX, ONNX, and also support browser-based inference using WebGPU.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.