RayRoPE: Ray Positional Encoding for Multi-View Attention
Key point
We propose RayRoPE, a positional encoding scheme with SE(3) invariance and geometric adaptability for multi-view transformers.
Details
Multi-view Transformers process tokens from a sequence of posed input images. Existing absolute or relative positional encoding methods have limitations in uniquely encoding patches while implementing SE(3)-invariant attention and adapting to the scene's geometric structure.
To address this, the proposed RayRoPE represents patch positions based on the associated Ray. In particular, instead of simply using direction alone, it fills the gap left by existing methods by leveraging a predicted Point along the ray.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.