AI Briefing
KO
Pick

Apple Releases LensVLM-9B for Processing Compressed Text Images

·2026.09.24 05:35

Key point

Based on Qwen3.5-9B, LensVLM maintains original-level accuracy at a 4.3x compression rate and demonstrates performance up to 10.1x.

Details

Apple researchers have released LensVLM to overcome the limitations of Vision-Language Models (VLMs) in processing text rendered as images. Existing VLMs suffer from a sharp drop in accuracy as compression rates increase when converting text to images, because characters become indistinguishable due to the resolution limits of visual encoders.

LensVLM applies an inference framework and post-training techniques that scan compressed images and selectively upscale relevant parts to original resolution using learned tools. Developed based on Qwen3.5-9B-Base, this model maintains accuracy similar to the original text upper bound at an effective compression rate of 4.3x, and outperforms existing retrieval-based and visual compression baselines up to a 10.1x compression rate across seven text QA benchmarks.

This approach enhances the robustness of visual compression regarding rendering choices and guides the model to rely on upscaled content rather than unreliable visual reading as compression rates increase. The analysis also provides practical guidelines indicating that text upscaling is more suitable for rendered text, while high-resolution image upscaling is better for native documents where layout information is important.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.