AI Briefing
KO

HuggingFace unveils 2B compact VLM 'SmolVLM'

·2024.11.26 09:00

Key point

HuggingFace has released SmolVLM, a 2B-scale open-source vision language model that maximizes memory efficiency.

Details

SmolVLM, developed by HuggingFaceTB, is a compact vision language model (VLM) family with 2B parameters. It is designed for deployment in local environments, browsers, and edge devices, offering excellent memory efficiency and inference speed.

The model is provided in three versions depending on use case:

  • SmolVLM-Base: A base model for downstream fine-tuning.
  • SmolVLM-Synthetic: A model fine-tuned on synthetic data.
  • SmolVLM-Instruct: An instruction-following model ready to use for conversational applications.

All model checkpoints, datasets, training recipes, and tools are released under the Apache 2.0 license, allowing commercial use. It was trained using open-source datasets such as The Cauldron and Docmatix, and is integrated into the Transformers library for immediate use.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.