Tencent Releases Youtu-Parsing-Omni: A 5B Omni-Modal Parsing Model for Unified Document and Media Analysis
Key point
The 5B-parameter model achieves state-of-the-art results on OmniDocBench v1.6 with a score of 96.96 and supports seven distinct parsing families via a unified JSON schema.
Details
Tencent has released Youtu-Parsing-Omni, a compact 5B parameter omni-modal parsing model designed to convert diverse inputs into a single structured JSON envelope. The model handles seven distinct parsing families, including document pages, natural images, charts, flowcharts, geometry figures, audio clips, and audio-visual videos.
Unified Schema and Capabilities
The model produces a single JSON output that covers both perception and cognition tasks. Perception includes layout elements, text, tables, formulas, bounding boxes, timestamps, ASR, OCR, acoustic events, and camera motion. Cognition tasks include captions, narratives, and reports. The specific output family is selected via a task prompt (e.g., --task in examples).
Supported modalities and key contents include:
- Document page: Layout elements with bbox, text/LaTeX/OTSL, tables, Mermaid flowcharts, and reading order.
- Natural image: Entities and text with bbox, tags, captions, and global description.
- Chart: Markdown table, notes, and caption.
- Flowchart: Mermaid code and caption.
- Geometry figure: Geometric relations, measurements, and elements.
- Audio: Vocal/non-vocal segments with timestamps, speakers, ASR, timbre/scene captions, and acoustic events.
- Natural video: Temporal segments with visual elements, actions, interactions, camera motion, and audio track.
- Text-rich video: OCR + ASR segments and a Markdown structured report.
Performance and Deployment
Youtu-Parsing-Omni demonstrates strong performance relative to its size:
- Achieved 96.96 Overall on OmniDocBench v1.6, marking a state-of-the-art result.
- Scored 75.08 Avg. on OmniParsingBench, ranking as the best open-weight model and second only to Gemini-3-Pro.
- Competitive with specialized models on chemical-structure (ChemOCR) and music-score (PDMX-Synth) recognition.
The release includes a vLLM plugin, pinned serving settings, task prompts, and inference examples to facilitate easy deployment.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.