Impact of Evaluation Resolution on V1 Similarity Assessment
Key point
It has been revealed that differences in evaluation resolution are the primary cause of the phenomenon where untrained CNNs exhibit higher similarity in V1 than trained CNNs.
Details
According to the arXiv preprint (2608.12408), it has been demonstrated that the claim that untrained CNNs exhibit higher representational similarity analysis (RSA) scores in the primary visual cortex (V1) than CNNs trained via backpropagation is primarily an artefact stemming from differences in evaluation resolution.
The research team used small CNNs trained on a CIFAR-10 subset at 32px resolution and evaluated THINGS-fMRI stimuli at six resolutions ranging from 32px to 224px. As a result, the V1 similarity gap between trained backpropagation (BP) models and untrained models showed a non-monotonic trend depending on resolution.
Specifically, the gap was negligible at -0.001±0.007 at 32px, but widened to +0.044±0.006 at 224px, where trained models exhibited higher similarity. This pattern was consistently observed across all 5 seeds.
Furthermore, commercial models trained at 224px, such as ResNet-50 and Swin-Tiny, also showed the same trend of peaking at low resolutions, confirming that the mismatch between training and evaluation resolution does not lead to artefactual results. Other factors, such as Gabor filter structure, uncorrected batch normalization, and global brightness convergence, were also ruled out.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.