VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models
Key point
**VSAS-Bench**, a new framework for evaluating the performance of real-time visual streaming assistant models, has been released.
Details
Existing VLM (Vision-Language Model) frameworks have primarily evaluated models in offline settings. However, streaming VLMs, which are core to real-time visual assistants, require not just simple video understanding but also proactiveness, which indicates the timeliness of responses, and consistency, which refers to the robustness of responses over time.
To address this, VSAS-Bench is proposed as a new framework and benchmark for visual streaming assistants. This benchmark has the following features:
- It provides over 18,000 temporally dense annotations across diverse domains and task types.
- It introduces standardized synchronous and asynchronous evaluation protocols.
- It provides metrics that can separately measure the individual capabilities of streaming VLMs.
Using this, the research team analyzed the accuracy–latency trade-off according to key design factors such as memory buffer length, memory access policy, and input resolution. Experimental results demonstrated that simply adapting existing VLMs to streaming settings without additional training can outperform state-of-the-art streaming-specific models. As one example, Qwen3-VL-4B achieved 3% higher performance than Dispider, the existing best-performing streaming VLM, under the asynchronous protocol.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.