Decomposing deep neural network-brain representational similarity reveals distinct sources across the human visual hierarchy
Abstract
Deep neural networks predict neural responses across the visual hierarchy, yet published alignment scores do not reveal whether this correspondence reflects learned representations, architectural inductive biases, low-level image statistics, or categorical structure. We decompose DNN-brain alignment into four sources using variance partitioning on representational similarity analysis, applied to 7T fMRI from eight participants viewing approximately 10,000 natural images. The composition of alignment shifts systematically along the cortical hierarchy: in primary visual cortex, architecture and low-level statistics account for 46% of total explained variance, whereas in high-level visual cortex model-specific learned representations dominate, reaching 68% in the parahippocampal place area. The same framework applied to language models processing image captions reveals that alignment is strongly model-dependent: a sentence embedding model contributes 51% unique variance in higher visual cortex, whereas an autoregressive model contributes only 4%, with the remainder attributable to surface text statistics and category structure. Across nine models, alignment scores varied substantially in composition. Decomposing alignment into its sources is necessary before a score can be interpreted as evidence for convergent processing.