LensVLM-9B: VLM that compresses text into images to surpass context limits
apple/LensVLM-9B
About the project
This is a 9B parameter model applying the 'visual-text-compression' technique, which converts text into images for compression. By encoding and processing long documents as images, it bypasses the context window limitations of existing VLMs.
Fine-tuned from Qwen3.5-9B, it features a multimodal architecture that accepts simultaneous image and text inputs to generate responses. It supports long-context capabilities for handling extended contexts and interactive dialogue.
Released under Apple's AMLR license, it can be easily deployed via the Transformers library. It is suitable for RAG and document analysis tasks requiring efficient processing of large-scale text data in image form.
apple/LensVLM-9B
The original page has no description.
image-text-to-text
This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.