AI Briefing
KO

DeepSeek-V4-Flash-Vision-Exp Released: Multimodal API Officially Open

·2026.08.21 09:00

Key point

DeepSeek has released V4-Flash-Vision-Exp, significantly enhancing multimodal agent performance while maintaining text capabilities.

1 / 2

Details

DeepSeek-V4-Flash-Vision-Exp has been officially released on the DeepSeek API platform. This experimental multimodal model adds visual understanding capabilities while retaining the text processing abilities (agents, reasoning, world knowledge) of the existing DeepSeek-V4-Flash.

It showed a significant performance improvement over V4-Flash in multimodal agent benchmarks, achieving multimodal agent performance close to Opus-4.8. It can be called with model='deepseek-v4-flash-vision-exp' and is supported by default starting from DeepSeek Harness 0.1.1.

Key API features and characteristics are as follows:

  • Multimodal Support: Supports mixed text and image inputs, billed at a maximum of 384 tokens per image.
  • Files API Introduction: Images can be uploaded once and reused by referencing file_id, saving request bandwidth.
  • Compatibility: Supports Chat Completions, Messages, and Responses APIs.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.