AI Briefing
KO

LFM2.5-VL-3B Enhances Edge Vision

·2026.08.12 23:00

Key point

Liquid AI released a 3.1B vision-language model with enhanced screen understanding and tool use.

Details

Liquid AI released LFM2.5-VL-3B, a 3.1B parameter vision-language model designed for on-device execution. It is designed to answer directly without generating lengthy reasoning traces, aiming for fast responses in real-time applications.

Key improvements include:

  • Screen and UI understanding across various devices
  • Object detection and grounding based on natural language queries
  • Information comparison and understanding across multiple images
  • Function calling supporting both text and image contexts

The model combines the SigLIP2 400M NaFlex vision encoder with the pre-trained backbone of the LFM2.5-2.6B text model. It was pre-trained on approximately 34 trillion tokens, with vision data increased fourfold compared to the previous version. The tokenizer vocabulary was also expanded to 128K to support non-Latin scripts.

Post-training involved knowledge distillation using a large teacher model, followed by SFT and Antidoom training, and finally multi-reward reinforcement learning. In official comparative evaluations, the average score across all vision tasks was 69.4, higher than the previous LFM2-VL-3B's 57.2.

Notably, it recorded 78.7 on ScreenSpot-v2 for desktop, 81.2 for mobile, and 82.2 for web, with a RefCOCO grounding score of 87.9. In tool-use evaluations, it scored 59.5 on ToolSandbox and 32.5 on BFCL V4, highlighting the potential for agent use in small models.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.