These results show that nested neural contrasts do not identify represented content by themselves: predictor construction is part of the experimental design, and matched controls are required for representational claims.
Abstract
Foundation-model features are increasingly used to ask what information neural activity represents, often by comparing prediction gains between nested encoding models. We show that such multimodal contrasts can change sign when only the conditioning predictor is reconstructed. Using fMRI from the Natural Scenes Dataset, DINOv2 visual features, and MPNet embeddings of MS COCO captions and Localized Narratives, a caption-narrative contrast in the additional predictive contribution of vision favors narratives when one short caption is compared with a long narrative (+0.012/+0.015 in Places), but favors captions after approximate word-count matching (-0.031/-0.023). The shift occurs across every measured ROI in both subjects and is driven primarily by differences in language-only prediction. Comparable contrasts also survive removal of image-specific content-word identity in several ROIs. These results show that nested neural contrasts do not identify represented content by themselves: predictor construction is part of the experimental design, and matched controls are required for representational claims.
It is found that a common result here -- that untrained or locally trained networks rival or beat backpropagation at early visual cortex -- depends strongly on the resolution at which the network is evaluated.
Deep neural networks predict neural responses across the visual hierarchy, yet published alignment scores do not reveal whether this correspondence reflects learned representations, architectural inductive biases, low-level image statistics, or categorical structure. We decompose DNN-brain alignment into four sources...
Resting-state connectivity can predict task-evoked fMRI activation, but correspondence with an individual task map may partly reflect a shared population pattern. We evaluated the Variational Resting-state-to-Task Prediction TransformeR (VRPTR), a three-dimensional encoder-decoder combining a compressed Transformer bot...
D. Di Giovanni, D. L. Collins· bioRxiv· 0 citations
It is proposed that CNN training's classification bottleneck compresses brain-relevant information at depth, unlike transformers's self-attention and non-classification objectives, which reflect signal strength and persistence rather than distinct brain regions.
S. Baghel, Kshitij Dwivedi, Dinesh Singh et al.· 0 citations
Understanding how the brain parses actions and events from time-varying natural inputs is a central challenge in neuroscience. Recent work has used deep neural network (DNN) models to build stimulus-computable fMRI encoding models that predict single-voxel responses to complex natural videos. However, the majority of v...
Iishaan Inabathini, Margaret M. Henderson· 0 citations
Computational hypotheses about brain information processing can be expressed in neural network models. Neuroscientists have begun to compare such models in terms of their alignment with neural and behavioral data. The high parametric capacity of these models is essential to their ability to capture cognitive processes...
Tal Golan, H. Schütt, N. Kriegeskorte· Nature Reviews Neuroscience· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.