NVIDIA Unveils Audex-30B, an Integrated Audio-Text LLM
·2026.07.07 16:12
Key point
NVIDIA has released Audex-30B, a unified multimodal MoE-based model that processes text and audio simultaneously.
Details
NVIDIA has unveiled Audex-30B-A3B, a new unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B. The model uses a 30B-parameter MoE (Mixture-of-Experts) architecture, but achieves efficiency by using only 3B active parameters for actual computation.
Key Features and Capabilities:
- Multimodal Capability: Implemented by extending an audio encoder for speech and general audio input and a discrete audio token vocabulary for text/audio output.
- Wide Range of Audio Tasks: Supports audio understanding, speech recognition and translation, TTS (Text-to-Speech), audio generation, and Speech-to-Speech generation.
- Preserved Performance: Integrates audio capabilities while causing almost no degradation to the existing text-only model's reasoning, alignment, knowledge, long-context, and agentic capabilities.
- Technical Specifications: Supports a context length of up to 1M tokens, and supports both a Thinking mode using
<think>tags and a regular Instruct mode.
Alongside this, NVIDIA is also providing a smaller model, Nemotron-Labs-Audex-2B.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.