AI Briefing
KO

NuExtract3: 4B VLM for Document Extraction Released

·2026.05.25 22:14

Key point

NuExtract3, a VLM based on Qwen3.5-4B specialized for document data extraction and Markdown conversion, has been released.

Details

NuExtract3, released by Numind, is a 4B-parameter open-weight VLM based on Qwen3.5-4B. It is optimized for extracting information from complex documents (PDFs, screenshots, forms, tables, receipts, etc.) and converting it into Markdown or JSON format.

Key features are as follows:

  • Document structure understanding: Effectively handles tables, forms, and pages with complex layouts.
  • Support for diverse inputs: Can process both text and visual document inputs.
  • Efficient local deployment: Can run with under 4GB of VRAM, and supports various quantization formats such as GGUF, MLX, GPTQ, and FP8.
  • Long-context processing: Trained for 3 days on an 8xH100 node, improving its ability to handle long documents.

It supports major inference engines such as vLLM, SGLang, and llama.cpp, and is freely available under the Apache-2.0 license.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.