StereoFoley: Video-Based Object-Aware Stereo Audio Generation
Key point
StereoFoley generated stereo audio that reflects object positions from video.
Details
StereoFoley is a video-to-audio framework that takes video as input and generates 48 kHz stereo audio. Existing models were strong in semantic alignment and temporal synchronization, but lacked stereo imaging that captures object-specific spatiality, and behind this was a shortage of professionally mixed spatial audio datasets.
The core consists of two stages.
- First, a base model was trained to achieve state-of-the-art semantic accuracy and synchronization.
- Next, a synthetic data pipeline combining video analysis, object tracking, and audio synthesis was built, using dynamic panning and distance-based loudness control to synthesize spatial audio matching object positions.
After fine-tuning with this synthetic dataset, the correspondence between objects and audio became sharper. For evaluation, a new stereo object-awareness metric was proposed and validated through listening experiments, and it was also confirmed that the metric correlates strongly with human perception. As a result, the first end-to-end framework for object-aware stereo video-to-audio generation was presented.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.