MiniMax-H3-Fun-Controlnet-Union: Control MiniMax-H3 with Five Control Videos in One
alibaba-pai/MiniMax-H3-Fun-Controlnet-Union
About the project
Process five types of control videos—Canny, Depth, HED, MLSD, and Pose—with a single checkpoint. Control the MiniMax-H3-based video generator on the VideoX-Fun pipeline without switching models per condition, as required by previous approaches. Video inpainting is also supported in the same manner.
The control branch connects to five of the 50 transformer blocks (layers 0, 10, 20, 30, and 40). With guidance distillation applied, setting guidance_scale to 1.0 enables inference in a single forward pass. Adjust the control_context_scale value to flexibly tune the influence of the control video from 0.0 to 1.0.
Output videos follow the frame count and aspect ratio of the control video, limited to a maximum length of 15 seconds. Rendering is fixed at 24fps, and prompts that detail the scene, subject, and camera movement contribute to stability. The checkpoint file contains only the control branch, approximately 6.8GB in size, and is loaded alongside the base MiniMax-H3 weights.
The transformer and text encoder combined require approximately 124GB of memory, making full loading difficult on a single 80GB GPU. Use the model_group_offload or model_cpu_offload_and_qfloat8 options for memory management. It is distributed under the MiniMax H3 Community License, and checking regional restrictions and usage policies is required.
alibaba-pai/MiniMax-H3-Fun-Controlnet-Union
The original page has no description.
text-to-video
This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.