AI Briefing
KO

Qwen3.5-LiveTranslate: From Sound to Sight, Translation Gets More Accurate

·2026.05.19 18:40

Key point

Qwen has unveiled Qwen3.5-LiveTranslate, a real-time interpretation model that leverages visual context as well.

Details

Qwen3.5-LiveTranslate-Flash is the latest simultaneous interpretation model built on Qwen3.5-Omni, which reads not only speech but also visual context together to improve translation accuracy. It targets multilingual environments such as international conferences, live streaming, online classes, and business negotiations.

The key upgrades are as follows.

  • Input audio/output text languages: 18 → 60
  • Output speech languages: 10 → 29
  • Average speech-to-speech per-token latency: 2.8 seconds
  • Voice cloning: pre-registered / clone-once / real-time 3 modes
  • Hotword: prioritized processing of up to 1,000 proper nouns and technical terms

Real-time performance was boosted through Readable Unit streaming. Compared to its predecessor Qwen3-LiveTranslate-Flash, first-token latency was reduced by 3.45 seconds and per-token latency by 1.88 seconds, with almost no degradation in translation quality, according to the company.

In offline evaluations on FLEURS and CoVoST2, it showed higher accuracy than mainstream commercial large speech models, and both language coverage and translation quality improved over the predecessor. The architecture uses a Thinker-Talker structure, where the Thinker produces text translations from audio and image inputs, and the Talker generates cross-lingual voice cloning from the original speech and the translated text.

The demos showcased international conferences, overseas travel, e-commerce livestreams, classical Chinese translation, and resolving visual ambiguity. The company stated that going forward it will expand toward lower latency, more languages and dialects, long-context consistency, more natural voice cloning, and interaction modes that encompass gestures, lip movements, and facial expressions.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.