MiniMax-H3: Simultaneous Generation of 2K Video and Stereo Audio with a Single 33B Stream
MiniMaxAI/MiniMax-H3
About the project
MiniMax H3 is a general-purpose generative system that holistically understands diverse multimodal inputs, including text, images, video, and audio. It follows complex multimodal instructions and generates native stereo audio alongside videos up to 15 seconds long at 2K resolution. The output frame rate is 24 FPS, and the audio sampling rate supports 32 kHz.
Two variants are provided based on the input method: H3-Base-FL2VA and H3-Base-Ref2VA. FL2VA accepts 0–2 images to perform text-to-video generation, first/last frame-based generation, or interpolation between the first and last frames. Ref2VA accepts up to 12 files as reference inputs, including up to 9 images, 3 video clips, and 3 audio clips, to handle more complex contexts.
The system consists of three modules: H3-Context-IR, H3-Base, and H3-Regenerate-2K. H3-Context-IR is a hosted preprocessing system that interprets free-form multimodal inputs and converts them into structured intermediate representations. H3-Base generates video and audio at 768p resolution based on these representations, while H3-Regenerate-2K reuses the low-resolution results and the original context to regenerate at 2K resolution.
H3-Base uses a single-stream Transformer architecture with 33B parameters, utilizing the pretrained weights of Qwen3-VL-32B as an encoder. Video and audio are encoded by H3-VisualVAE and H3-AudioVAE, respectively, and processed as an integrated multimodal sequence. The current open-source release includes the H3-Base weights, while H3-Context-IR and H3-Regenerate-2K can be accessed via API or by following the official workflow.
MiniMaxAI/MiniMax-H3
The original page has no description.
image-text-to-video
This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.


