AI Briefing
KO

LensVLM-9B: VLM that compresses text into images to surpass context limits

apple/LensVLM-9B

·2026.09.27 05:52

This is a 9B parameter model applying the 'visual-text-compression' technique, which converts text into images for compression. By encoding and processing long documents as images, it bypasses the context window limitations of existing VLMs.

Fine-tuned from Qwen3.5-9B, it features a multimodal architecture that accepts simultaneous image and text inputs to generate responses. It supports long-context capabilities for handling extended contexts and interactive dialogue.

Released under Apple's AMLR license, it can be easily deployed via the Transformers library. It is suitable for RAG and document analysis tasks requiring efficient processing of large-scale text data in image form.

HuggingFace
HuggingFace model

apple/LensVLM-9B

The original page has no description.

image-text-to-text

This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.