V-JEPA 2.1 Robustness: Not Proportional to Model Size
Key point
An analysis of Meta's V-JEPA 2.1 models confirmed a non-monotonic scaling property in which robustness does not necessarily improve as model size increases.
Details
A pre-registered robustness study was conducted on Meta's V-JEPA 2.1 models (80M–2B), confirming three key findings.
1. Partitioned Dense Features The model's dense features respond differently to temporal corruption and image-noise. Temporal corruptions such as frame drops or occlusion were good predictors of task failure, but correlation with image noise such as Gaussian noise or motion blur was not statistically significant.
2. Non-monotonic Relationship Between Model Size and Robustness Larger model size does not always lead to improved robustness. Non-monotonic results were observed across all Tier 1 corruptions, and notably, the 2B model showed lower robustness than the 1B model on three corruption types.
3. Orientation-sensitivity V-JEPA 2.1 is sensitive to orientation. When a horizontal flip is applied, the temporal structure is preserved, but the representation is disrupted in a manner similar to playing the video in reverse. This suggests that the model fundamentally lacks orientation-equivariance.
As a possible cause of this non-monotonic scaling, the phenomenon of hub marginalization occurring in deep ViTs has been proposed. The analysis suggests that as the model gets deeper, additional layers may enter a stage where they scramble information rather than refine it.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.