Dataset Fragmentation, Cognitive Variability, and Reproducibility Challenges in EEG-Based Brain–Computer Interfaces: A PRISMA-Based Systematic Review
Abstract
Highlights What are the main findings? EEG-BCI dataset fragmentation spans structural, semantic, procedural, computational, and human-contextual dimensions; common file access or statistical alignment alone does not establish dataset comparability. Within the curated 129-publication synthesis corpus, the median unweighted reporting-transparency score was 9/10, while code or pipeline availability was stated in only 20.9% and cognitive and user-state metadata were represented inconsistently. What are the implications of the main findings? Scalable cross-dataset learning requires analysis-dependent compatibility rules, explicit provenance, contextual metadata and auditable transformations, rather than relying on format conversion alone. Existing standards, software platforms, benchmark frameworks, ontologies and transfer-learning methods should be connected as complementary interoperability layers, with uncertainty and information loss reported explicitly. Abstract Public EEG-based brain–computer interface (BCI) datasets are expanding rapidly, yet differences in sensors, experimental protocols, task/event semantics, preprocessing, participant context, and evaluation limit reproducibility and cross-dataset learning. This review examined whether heterogeneous EEG-BCI resources can support reproducible analysis across sources. A PRISMA-based systematic mapping review of literature published from 2014 to June 2026 was conducted across Scopus, Web of Science Core Collection, IEEE Xplore, PubMed, and ACM Digital Library, yielding 16,920 records. Following screening and evidence-focused curation, a curated synthesis corpus of 129 publications was retained and confirmed by full-text review. Evidence was coded across structural, semantic, procedural, human/contextual, and computational fragmentation, and reporting transparency was assessed using ten criteria. The synthesis identified heterogeneity in acquisition, channel layouts, task/event definitions, preprocessing, participant/session context, and evaluation design. Existing standards, ontologies, software platforms, benchmark frameworks, and transfer-learning methods address complementary layers but do not provide complete semantic interoperability. Within the retained corpus, the median transparency score was 9/10; data availability was stated in 61.2% and code or pipeline availability in 20.9%. These frequencies describe the curated corpus rather than the field as a whole. Scalable cross-dataset analysis requires analysis-dependent compatibility rules, explicit provenance, contextual metadata, and auditable transformations that preserve dataset identity, uncertainty, and information loss.