AI Briefing
KO

Interaction-Aware Video Object Removal (VOID)

·2026.04.06 09:00

Key point

Netflix has unveiled VOID, a video inpainting model that naturally handles physical interactions as well when removing objects.

Details

VOID handles not only secondary effects like shadows or reflections when removing objects from video, but also the physical interactions the object has with its surrounding environment. For example, when removing a person holding a guitar, it also reproduces the physical change of the guitar naturally falling to the floor.

The model is built on CogVideoX, and improves video inpainting performance through interaction-aware mask conditioning.

The model sequentially uses two-stage transformer checkpoints to improve temporal consistency.

  • VOID Pass 1: the base inpainting model
  • VOID Pass 2: a refinement model using Warped-noise

For mask generation, it utilizes the Gemini API and SAM2, and 40GB or more of VRAM (e.g., an A100) is required for smooth inference.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.