PickFastVideo-FastH3-4-step-Preview-v1-VSA-DataFree: Simultaneous Video and Audio Generation with 4-Step Inference
FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree
About the project
Generate synchronized video and audio outputs from a single text prompt. A 35B parameter model based on MiniMax H3 is compressed using the DMD2 distillation technique, completing the generation process in just 4 Transformer forward passes.
Computational efficiency is improved by applying the VSA-H3 attention backend and 90% sparsity. While complex motion, fine details, and some audio quality may be lower compared to the original model, it is useful for rapid prototyping or in resource-constrained environments.
Default configurations tested on a 4x B200 GPU setup are provided. CUDA 13 and the Blackwell architecture are prioritized, while other multi-GPU environments can run via Triton kernels and specific flags.
FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree
The original page has no description.
text-to-video
This introduction was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report errors, attribution issues, or removal requests via Contact.