AI Briefing
KO

Unified Vision Model SenseNova-Vision Released

·2026.08.14 00:07

Key point

The 7B-scale open model SenseNova-Vision has been released, capable of performing various vision tasks without separate heads.

Details

SenseNova-Vision is a 7B-scale MoT (Mixture of Tasks) model that performs various computer vision tasks using only natural language instructions, without separate task-specific heads.

The core concept of this model is treating all vision tasks as a single generation problem. When users provide natural language commands or visual hints, the model can output both text and images.

Key Features and Characteristics:

  • Diverse Task Support: Performs object detection, keypoints, OCR, various segmentation types (Binary, Instance, Semantic), depth estimation, and surface normal estimation.
  • 3D and Camera Estimation: Equipped with multi-view 3D reconstruction and camera pose estimation capabilities, allowing implementation via prompts alone without specialized tools like COLMAP.
  • Training Data: Trained on 50 million instruction-response pairs built from diverse CV annotations.
  • License and Accessibility: Released under the Apache 2.0 license, with weights available on Hugging Face.

However, running the web demo smoothly requires a high-performance GPU with 80GB or more, and higher-spec hardware may be required to fully test the model's performance.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.