Skip to content
Open access

Reliability and disease sensitivity are dissociable properties of EEG foundation-model representations

Aug 2026 · bioRxiv · 0 citations · 30 references
Biology

TL;DR

Reliability showed no detectable association with discrimination, pretraining paradigm, or domain, and had to be measured directly, so it is recommended to become a standard evaluation axis for EEG-FM representations intended for longitudinal or biomarker use, and release a reproducible pipeline.

Abstract

EEG foundation models (EEG-FMs) are evaluated almost entirely on disease-discrimination accuracy. A clinical biomarker additionally requires measurement reliability, the stability of repeated measurements on the same individual, which regulatory biomarker frameworks treat as a prerequisite that discrimination does not imply. We asked whether frozen EEG-FM representations provide such stability, whether it is predictable from conventional model descriptors, and what information supports it. We measured test-retest reliability, disease discrimination, and representation distinctiveness for nine frozen representations: six EEG-oriented foundation models spanning masked, contrastive and predictive pretraining, handcrafted spectral features, and two general-purpose time-series models with no EEG exposure. All were evaluated under one preprocessing pipeline across two healthy retest cohorts, at roughly one month and two years, and three neurodegenerative cohorts. Reliability was measured in healthy adults only; the disease cohorts contribute cross-sectional discrimination. Reliability, quantified as the intraclass correlation coefficient (ICC), varied enormously (mean 0.08 to 0.76; coeffiicient of variation, CV, 53.0%) while disease discrimination, measured as area under the receiver operating characteristic curve (AUC), occupied a far narrower range across the same nine (AUC CV 5.4%), a roughly tenfold difference in relative dispersion, described rather than formally tested. We did not test formal AUC equivalence, so we describe discrimination as varying substantially less than reliability rather than as equivalent. The variation was not consistently explained by pretraining paradigm or domain among the models studied, and a model with no EEG exposure was among the most reliable tested. Alpha-band information contributed disproportionately to reliability, whereas theta-band information ranked first for Alzheimer’s disease and frontotemporal dementia discrimination, directionally consistent with established EEG evidence in both conditions. Only the reliability half of that contrast is individually significant. Subspace geometry and band ablation, two methodologically distinct analyses, both indicate that the two properties are partially, not fully, dissociable, and network architecture determines whether the dissociation is preserved, traded off, or jointly degraded across depth. Reliability showed no detectable association with discrimination, pretraining paradigm, or domain, and had to be measured directly. We recommend it become a standard evaluation axis for EEG-FM representations intended for longitudinal or biomarker use, and release a reproducible pipeline.

Read PDF

Similar papers

Open access Sep 2026

Quantifying Subject-Identity Variance in Spectral EEG Features and Its Role in Machine Learning Evaluation Leakage

Electroencephalography (EEG) is widely used to study cognitive states and clinical biomarkers, yet the contribution of stable between-subject differences to common EEG feature representations is rarely quantified directly or incorporated into evaluation design. Here, we examine subject-linked structure in spectral EEG...

Hassan Ugail, Richard Wirt, Newton Howard · 0 citations
Open access Sep 2026

Explainable machine learning for Alzheimer's disease characterization using small-sample EEG data

Alzheimer's disease (AD) is associated with progressive cognitive decline and altered brain functional activity, yet objective and interpretable electrophysiological indicators remain insufficiently established. Resting-state electroencephalography (EEG) offers a low-cost and clinically accessible candidate, provided t...

Lang Shen, Wei Tong, Ye Zhao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EEG-Xplain: Decoding Neural Black-Boxes of EEG Foundation Models

A unified attribution framework for interpreting EEG foundation models across heterogeneous architectures is proposed, providing a standardized approach for evaluating the interpretability, reliability, and physiological plausibility of EEG foundation models.

Han-Song Ma, Jun-Xiao Wang · 1 citation
Preprint Sep 2026

An Evidence-Aware Framework for EEG Microstate Analysis: Improved Sensitivity to Alzheimer's Disease and Ageing

Electroencephalography (EEG) microstate analysis commonly converts each scalp topography into a winner-take-all hard label and summarises the resulting sequence using duration, occurrence, coverage, transitions, and symbolic complexity. Although interpretable, this readout discards evidence strength, assignment ambigui...

Kai-Dong Wu, Hai-Li Ye, P. Sarrigiannis et al. · 0 citations
#machine learning Preprint Sep 2026

Representational and Functional Robustness to Electrode Montages in EEG Foundation Models

This work investigates the effects of different electrode montages through a joint functional and representational analysis of four EEG foundation models selected to span distinct montage-handling designs, and shows that aggregation, not the encoder alone, determines functional robustness.

Jakob Steglich, Justus Meyer zu Bexten, Shakiba Moradi et al. · 0 citations
Preprint Aug 2026

EEG-PRISM: Physiologically-Grounded Interpretability of Predictions by EEG Foundation Models

Objective: Foundation models represent the next advancement in AI for EEG analysis; however current explainable AI techniques provide attribution scores in the time-channel input space, which is mismatched to clinical intuition about EEG. Thus, there is a critical need for a universal method that can extend the interpr...

Deeksha M. Shama, Punnisa Amornsirikul, Archana Venkataraman · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.