Text-Conditioned JEPA for Learning Semantically Rich Visual Representations
Key point
The paper proposes TC-JEPA, which learns semantically rich visual representations by conditioning on captions.
Details
The existing I-JEPA predicts latent features at masked locations, but semantic learning was difficult due to the uncertainty of occluded regions.
TC-JEPA conditions on image captions and adjusts the predicted patch features via sparse cross-attention over the input text tokens. As a result, patch representations become better predicted by the text, leading to more semantic visual representations being learned.
- downstream performance, training stability, and scalability all improved or showed promise.
- It presents a new vision-language pretraining paradigm that trains solely through feature prediction.
- It outperformed contrastive methods, especially on tasks requiring fine-grained visual understanding and reasoning.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.