Skip to content

BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings

Sep 2026 · arXiv.org · 0 citations
Computer Science Biology Medicine

TL;DR

BrainWideBench is presented, a benchmark for evaluating across-animal transfer on multi-region neural recordings, built on the International Brain Laboratory Brainwide Map dataset of neural and behavioral recordings spanning 276 brain regions from 139 mice performing a sensory-guided decision-making task.

Abstract

Advances in large-scale neural recording have made it possible to collect data across many animals and distributed brain regions, raising the question of whether this scale can be exploited to learn general-purpose neural representations transferable across diverse downstream tasks. Yet, progress toward this goal has been limited by fragmented evaluation protocols and a narrow focus on individual task domains. Here, we present BrainWideBench, a benchmark for evaluating across-animal transfer on multi-region neural recordings, built on the International Brain Laboratory Brainwide Map dataset of neural and behavioral recordings spanning 276 brain regions from 139 mice performing a sensory-guided decision-making task. The benchmark is organized around three complementary task suites that evaluate whether learned representations support downstream decoding of behavior, can predict masked or future neural activity, and can recover biologically meaningful anatomical organization. With this benchmark, we systematically evaluate pretraining methods across transfer settings, including finetuning on downstream objectives and zero-shot generalization to unseen animals. Our results confirm pretraining improves performance over matched single-session baselines, but we show current methods exhibit heterogeneity in transfer capabilities: gains depend strongly on the alignment between pretraining objectives and downstream tasks. No single approach performs uniformly well across all three suites, and most methods are designed to only address a subset of them. Together, these findings suggest that learning representations that jointly generalize across behavior, dynamics, and anatomy remains an open challenge. By providing a unified and reproducible evaluation suite, BrainWideBench establishes a framework for measuring progress toward general-purpose models of the mouse brain. Code is available at this link.

View source

Similar papers

#machine learning Preprint Sep 2026

Pretraining for Sample-Efficient Neural Interfaces

This work proposes MAPA, an otherwise vanilla masked autoencoder with two spatial encodings, an anatomical region embedding and a relative positional encoding that together enable it to learn neural representations that transfer to unseen subjects and across various tasks.

Ben-Ting Tang, Z. Spalding, G. Cogan · 0 citations
#machine learning Preprint Sep 2026

iMINDBench: iEEG Multi-Institution Neural Decoding Benchmark

Intracranial electroencephalography (iEEG) is widely used to record electrical activity directly from electrodes inside the human brain, making it an attractive modality for neural decoding. However, progress in iEEG decoding, especially toward general-purpose foundation models, remains difficult to measure reliably: d...

Geeling Chau, Saba Hashemi, Yonghyeon Gwon et al. · 0 citations
Preprint Sep 2026

BrainTaskonomy: Learning How to Pretrain and What to Transfer in fMRI Foundation Models

fMRI foundation models increasingly aggregate heterogeneous data across brain states, cohorts, and acquisition settings, yet pretraining domains are commonly treated as a flat mixture and downstream tasks are adapted independently. We study whether measured learning relations can organize both stages without modifying...

Jun-Feng Xia, Wen-Hao Ye, Jun-Xiang Zhang et al. · 1 citation
Open access Sep 2026

NeuroGate: waveform translation between brain regions

The brain processes information across distributed circuits, yet a typical experiment records only a few regions, leaving the rest unobserved. Connectivity and latent-embedding methods relate brain regions but do not return the waveform of an unrecorded one. Here we introduce NeuroGate, a framework for cross-regional n...

Ali Zareh, Rufeyda Yağcı, M. K. Özdemir · 0 citations
Preprint Aug 2026

NeuroPB: Scaling Neural Decoding with Pretrained Behavioral Representations

Decoding continuous motor trajectories from neural activity is essential for developing practical brain-computer interfaces (BCIs). However, current neural decoders are constrained by the limited scale and heterogeneity of neural recordings. In contrast, behavioral data can be collected more readily and at substantiall...

Lu-Yao Jin, Yonghao Song, Huan Zhao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

iBrain: A Unified Foundation Model Reading the Brain from Surface to Spikes

Invasive neural recordings provide high-fidelity measurements of brain activity, with signals such as intracranial EEG (iEEG) and intracortical spiking activity capturing neural dynamics at different spatial and temporal scales. Yet existing neural foundation models have largely been developed independently for differe...

Ying Chen, Ti-Ou Wang, Zhi-Feng Yue · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.