FireRedAudio: 9B Unified Audio LLM Released
Key point
FireRedTeam released FireRedAudio, a 9B parameter unified audio language model that decouples understanding and generation.
Details
FireRedTeam has released FireRedAudio and FireRedTTS3 on Hugging Face and GitHub. Based on a 9B parameter LLM, FireRedAudio features a design that decouples the representations of the audio encoder (understanding) and the RedAE pathway (generation), performing ASR, audio understanding, zero-shot TTS, instruction-based TTS, and voice editing in a single model.
It enables accurate temporal grounding and reasoning for 1 hour of audio, demonstrating competitive performance on major benchmarks such as MMAU, MMSU, and Seed-TTS-Eval. FireRedTTS3 is a unified speech generation and editing model that utilizes semantically enriched speech representations.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.