AI Briefing
KO

DeepSeek-V4-Flash-Vision-Exp: Experimental Multimodal Model Adding Visual Understanding While Maintaining Text Agent Performance

deepseek-ai/DeepSeek-V4-Flash-Vision-Exp

·2026.08.31 23:16

This is an experimental multimodal model that combines a visual module with the DeepSeek-V4 Flash architecture to process text and images simultaneously. It significantly improves scores on multimodal agent benchmarks such as ApexBench while maintaining existing text-only agent performance.

It recorded 83.9 points on Terminal Bench 2.1 and 59.3 points on DeepSWE, indicating that its text-based coding and terminal task capabilities are close to Opus-4.8. Its ApexBench Pass@1 score is 36.5, an improvement of more than 10 points compared to previous versions that ignored multimodal inputs.

It provides an inference implementation reflecting the latest architectures such as DFlash attention, MoE, and Hyper-Connections. Enabling the DSPARK inference algorithm in SGLang allows for efficient serving by sharing target and draft weights within a single checkpoint.

Released under the MIT license, it allows for commercial use, but due to its strongly experimental nature, validation is required before production deployment. It is a suitable option for developers who need to add visual recognition without sacrificing text agent performance.

HuggingFace
HuggingFace model

deepseek-ai/DeepSeek-V4-Flash-Vision-Exp

The original page has no description.

image-text-to-text

This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.