Qwen2-Audio: Talking with Your Voice
Key point
Qwen's next-generation audio model, **Qwen2-Audio**, has been released, supporting both voice and text input.
Details
Multimodal understanding capability is essential for building AGI systems. Following Qwen-VL and Qwen-Audio, which extended the existing Qwen model to vision and audio, the next-generation model Qwen2-Audio has been released.
Qwen2-Audio is a model that takes audio and text input and responds with text, offering the following key features.
- Voice Chat: Understands users' voice commands directly without a separate ASR (Automatic Speech Recognition) module.
- Audio Analysis: Analyzes various audio information such as speech, sounds, and music according to text instructions.
- Multilingual: Supports 8 or more languages and dialects, including Chinese, English, Cantonese, French, Italian, Spanish, German, and Japanese.
Currently, the Qwen2-Audio-7B and Qwen2-Audio-7B-Instruct models have been released as open-weight on Hugging Face and ModelScope. These models demonstrate excellent performance across a wide range of tasks, including speaker identification, speech translation, background noise detection, and audio-based storytelling.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.