AI Briefing
KO

vdn-minimax-h3: 11 Seconds to Generate a 14-Second Video: Maximizing Speed with Hybrid Attention

OpenVDN/vdn-minimax-h3

·2026.09.04 02:23

A hybrid attention architecture based on MiniMax H3 has significantly improved video generation speed. In an environment with 8 B200 GPUs, it generates a 14.4-second video in just 11.23 seconds, achieving an inference speed faster than real-time playback.

The core of the architecture is the combination of frame-level linear attention and Softmax attention. It improves computational efficiency while maintaining the visual quality and consistency of the existing backbone, implemented in a plug-and-play manner by adding a separate linear attention branch and LoRA adapters.

Both the accelerated inference stack and training code are open-sourced, making them immediately usable in research and development environments. It supports FP8 quantization and achieves overwhelming speed improvements over 50-step models with 8-step denoising.

However, attention must be paid to the scope of the license. Under the MiniMax H3 Community License, usage may be restricted in specific regions such as South Korea, the United States, and Europe, so checking the terms before deployment is essential.

HuggingFace
HuggingFace model

OpenVDN/vdn-minimax-h3

The original page has no description.

text-to-video

This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.