Controlling Text-to-Audio-Synchronized Video Generation via Advanced Modality Conditioning and Interaction
Key point
The research presents a solution to the cross-modal interference and caption inconsistency problems that arise when generating audio-video synchronized footage from text.
Details
Text-to-Sounding-Video generation (T2SV) is a technology that creates footage in which audio and video are aligned according to text conditions. However, existing approaches face two key challenges.
The first is the bottleneck of Text Conditioning. The shared captions used during training cause cross-modal interference, and there is a gap between the detailed captions used for training and the concise prompts actually used at inference time. The second is that the optimal fusion mechanism for cross-modality feature interaction remains unclear.
To address this, we propose the CRR(Cross-Referential Rewriter) captioning framework. This framework is a pipeline composed of two agents.
- Semantic Checker: extracts well-grounded Semantic Anchors.
- Cross-Modal Rewriter: generates a disentangled pair of captions (TV and TA).
Through this approach, cross-modal interference is eliminated and the gap between training and inference is resolved.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.