AI Briefing
KO
Pick

MiniMax Music 3: A 5-Minute Music Generation Model Released

·2026.08.16 20:30

Key point

The MiniMax Music 3 model has been released, capable of generating consistent music up to 5 minutes long based on lyrics and music descriptions.

Details

MiniMax Music 3 is a music generation model that adopts a hierarchical structure to simultaneously solve the challenges of maintaining long-term structure and fine-grained acoustic quality in music generation.

Core Architecture: Hybrid LM

  • Global LLM (8B): Models the long-term structure and semantic flow of the song, predicting the first RVQ codebook.
  • Local LLM (0.6B): Predicts the remaining codebooks to reconstruct fine-grained acoustic information within each frame.
  • Synthesis Stage: Instead of discrete tokens, it fuses the continuous hidden states of the global and local models to generate high-quality audio via Flow Matching and Flow-VAE.

Control and Input Methods Users can achieve precise control through Structured Captions along with lyrics.

  • Global Metadata: Genre, BPM, key, emotional progression, etc.
  • Vocal Details: Vocal gender, timbre, style, etc.
  • Arrangement: Instrumentation, section-specific changes, texture, etc. Additionally, section tags such as [Verse] and [Chorus] can be used to explicitly define the song's structure.

Deployment and License Information

  • Hardware Requirements: 2 CUDA GPUs are required for inference, and only non-streaming mode is currently supported.
  • Serving: Can be served via SGLang-Omni.
  • License: Follows the MiniMax-Music3 Community License. Commercial use requires attribution of the model name, and separate permission is required if annual revenue exceeds $20 million.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.