Skip to content
Open access

Dynamic sample augmentation discovers gene biomarker for IgA nephropathy

Aug 2026 · Journal of Biological Engineering · 0 citations

TL;DR

A dynamic data augmentation method based on Layer-wise Relevance Propagation (LRP) aimed at improving classification performance and biological interpretability under small-sample conditions and through this open-environment enrichment approach, six hub genes were identified as gene markers for IgAN.

Abstract

Bulk RNA-seq data suffers from the issues of “high dimensionality and small sample size,” which limits its application in disease research. This paper proposes a dynamic data augmentation method based on Layer-wise Relevance Propagation (LRP) aimed at improving classification performance and biological interpretability under small-sample conditions. The method utilizes the LRP algorithm to calculate the contribution weight of each gene to the classification result and uses this weight to guide sample generation. By systematically amplifying biologically meaningful signals, it constructs semantically reliable augmented samples, avoiding the semantic distortion caused by traditional random perturbations. Simultaneously, a dynamic augmentation mechanism is introduced that tightly couples sample generation with model training, providing difficult-to-classify samples with multiple iterative optimization opportunities and forming a virtuous cycle where classification performance and augmentation quality improve synergistically. On this basis, population-level gene biomarkers are identified from the trained model. Innovatively, an open-environment enrichment analysis method is proposed—that is, instead of being limited to a few feature genes of a single subtype, the union of feature genes from all disease subtypes is taken for enrichment analysis, revealing shared biological pathways from a systems-level perspective and providing a more comprehensive interpretation for subtype-specific mechanism research. Experimental results show that this method effectively improves classification accuracy, and through this open-environment enrichment approach, six hub genes were identified as gene markers for IgAN.

Read PDF

Similar papers

Open access Jul 2026

Explainable Artificial Intelligence for Cross-Dataset Generalizable Biomarker Discovery in Cardiovascular diseases (CVDs)

CVDs are heterogeneous, multifactorial disorders that remain the leading cause of global mortality from infancy to old age. It requires an early identification and treatment of risk factors to accelerate disease prevention and morbidity improvement. Advancements in transcriptomics technologies gives large pool of heterogenous gene expression data. The technical heterogeneity of gene expression data reduces ability to compare multiple cross-platform datasets at once. To bridge gap, we systematically evaluate three data harmonization techniques: Shambhala-2, TDM, and UPC to align heterogeneous data into a shared expression space while preserving biological signals. Our pipeline integrates 25 independent datasets comprising 983 samples across 23 distinct CVDs phenotypes from both RNA-seq and microarray platforms. The framework benchmarks 35 Machine learning (ML) and Deep learning (DL) classifiers, including Transformers and ResNets, across three data modalities such as RNA-seq, microarray hybridization and RNA-seq + microarray and multiple tissue types. To ensure clinical trustworthiness, we apply multiple Explainable artificial intelligence (XAI) methods, such as SHapley additive exPlanations (SHAP) and Integrated gradientss (IGs), and assess their reliability using quantitative metrics like Area over the perturbation curve (AOPC), Sensitivity, and Infidelity. Results indicate that Shambhala-2 provides superior harmonization by maximizing the biological signal-to-platform ratio. Evaluation of XAI methods reveals that Shapley-based approaches offer the highest stability for identifying influential genomic features in high-dimensional data. Functional enrichment and pathway analyses further confirmed the involvement of identified biomarkers in key cardiovascular processes, including inflammation, immune regulation, oxidative stress, and vascular remodeling. Collectively, this study provides a scalable and interpretable road-map that integrates XAI with cross-dataset biomarker discovery, supporting the transition toward precision cardiology.

Ahtisham Fazeel Abbasi, Muhammad Sajjad, Sebastian J. Vollmer et al. · 0 citations
Open access Sep 2026

Dimensionality reduction and metabolite panel derivation in urinary metabolomics based on Random Forest with Gini index feature selection.

Untargeted urinary metabolomics represents a promising approach for investigating metabolic alterations associated with oncogenic processes such as breast cancer (BC). However, the stable selection of informative m/z features remains a central challenge in biomarker-oriented studies, particularly in the context of early BC detection and population-level screening. Urine samples from two independent cohorts (n = 50 and n = 75) and analyzed using distinct UHPLC-QTOF-ESI⁺ mass spectrometry workflows, yielded 224 and 129 aligned m/z features, respectively. Classification and embedded feature selection were implemented within a leakage-controlled Random Forest (RF) framework using Gini index-based importance ranking. Model evaluation incorporated repeated train-test splits and cross-validation to ensure methodological rigor and minimize overfitting. By consistently applying the same RF-based analytical framework to two analytically distinct cohorts generated under different chromatographic separation conditions, we demonstrate that a unified supervised strategy can achieve comparably high classification performance despite differences in feature dimensionality. Further, controlled dimensionality reduction identified compact panels of 25 m/z features per cohort while preserving classification performance. Importantly, stability was maintained after feature reduction, with strong accuracy, F1 scores, and receiver operating characteristic and precision-recall characteristics observed in both full and reduced models. This cross-cohort consistency indicates that the discriminative signal captured by the RF approach is not cohort-specific nor dependent on a particular separation workflow, but rather reflects reproducible metabolic patterns associated with BC. The stability of feature selection was further supported by substantial overlap between RF-derived Gini importance rankings and variable importance in projection (VIP) scores obtained from partial least squares discriminant analysis (PLS-DA) in MetaboAnalyst 5.0, indicating concordance across distinct supervised multivariate frameworks. Collectively, these findings highlight the advantage of a unified, supervised tree-based strategy capable of delivering stable classification and interpretable dimensionality reduction across independent untargeted metabolomics platforms, providing a structured and transferable framework for metabolomics-driven biomarker discovery and future clinical validation.

Markus Zetes, Vlad Moisoiu, Carmen Socaciu et al. · 0 citations
Open access Aug 2026

A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings

This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention and identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets.

Erik Vonkaenel, Lisa M. Bramer, J. Flores et al. · 0 citations
Open access Sep 2026

A Disease-Guided Representative Gene Selection Framework for High-Dimensional Gene Expression Analysis

Gene expression datasets provide valuable information for disease classification and biomarker discovery; however, their high dimensionality and limited sample size may limit classification performance and reduce biological interpretability. This study proposes GeDiRep, a prior knowledge-guided framework for identifying compact and informative gene subsets. The method first organizes filtered genes into disease-associated groups using curated gene–disease associations from DisGeNET. Each group is then scored according to its predictive contribution, and representative genes are selected using Random Forest-based feature importance. Representative genes from the top-ranked groups are progressively accumulated, and the resulting gene subsets are used to assess classification performance on the test set. Experiments on eight microarray datasets showed that GeDiRep reduced the average number of selected features from 40.3 to 8.3 compared with G-S-M while improving the average AUC from 0.83 to 0.87. In comparison with traditional feature selection methods using the same number of genes, GeDiRep also achieved competitive AUC values. Biological analyses, including term–gene network, hub gene, and heatmap analyses, supported the functional relevance and stability of several selected genes. Overall, GeDiRep provides a structured and interpretable framework for high-dimensional gene expression analysis by selecting reduced yet discriminative and biologically meaningful gene subsets.

Cihan Kuzudisli, B. Qaqish, Burcu Bakir-Gungor et al. · 0 citations
Open access Sep 2026

Adversarial random forests for omics synthesis

Data availability is critical for understanding complex disease pathways and developing robust predictive models. Although high-throughput omics technologies have improved insight into disease mechanisms, data acquisition from inaccessible tissues such as the central nervous system remains a major limitation, causing small sample sizes and complicating early prediction of neurodegenerative disorders such as Alzheimer’s and Parkinson’s diseases. Generative modeling has emerged as a powerful approach for synthesizing data to support downstream clustering and prediction with small sample size, but existing methods rarely handle high-dimensional tabular omics data effectively. Adversarial random forests (ARFs) provide a well-performing framework for tabular data generation but are not designed for high-dimensional settings. To address this limitation, we introduce high-dimensional ARF (h-ARF), an extension of ARF optimized for integrated clinical and high-dimensional omics data. Using benchmarks across nine datasets and eight performance metrics, we show that h-ARF better preserves both feature distributions, and downstream clustering and prediction utilities compared with ARFs. The method is implemented in the opensource R package harf, available on CRAN.

C. Fouodo, J. Kapar, Anke Huels et al. · 0 citations
Review Open access Aug 2026

AI-Driven Multi-Omics Integration for Early Disease Detection: A Comprehensive Survey

By comparing genomic and multi-omics strategies, this survey highlights ongoing hurdles related to privacy, bias, interpretability, and scalability.

Ahamadi Firdose, Deepthi Raj D, Bhargavi B, Jenita J, Apoorva H G, Dr. Madhu Gopinath · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.