Skip to content
Open access

stably: error-controlled stability selection for biomarker panel discovery

Sep 2026 · bioRxiv · 0 citations · 1 references
Biology

TL;DR

Stably, a Python package that implements stability selection with formal false positive control for DIA proteomics data, is developed and shows that the Shah and Samworth complementary pairs stability selection framework recovers more true synthetic biomarkers than the Meinshausen and Bühlmann framework at moderate effect sizes typical of proteomics.

Abstract

Motivation Data-independent acquisition mass spectrometry (DIA-MS) has become increasingly popular for clinical proteomics due to its sensitivity and reproducibility. Univariate statistical analysis tools, such as limma and MSstats, are widely used for identifying differentially abundant proteins but cannot capture multivariate relationships between proteins that may provide greater discriminatory power as a panel. Machine learning approaches can address this gap, but typically prioritise predictive performance over feature stability, producing biomarker panels for downstream validation that vary depending on data splitting and are poorly suited to clinical translation. Results We developed stably, a Python package that implements stability selection with formal false positive control for DIA proteomics data. Using synthetic data with known ground-truth biomarkers, we show that the Shah and Samworth complementary pairs stability selection framework recovers more true synthetic biomarkers than the Meinshausen and Bühlmann framework at moderate effect sizes typical of proteomics (d = 0.5 – 2.0), while both maintain false positive rates well below their theoretical guarantees. Applied to a publicly available serum proteomics dataset from patients with all stages of pancreatic ductal adenocarcinoma (n=176), stably identified a stable 17-protein biomarker panel in the discovery cohort (n=120), which achieved higher predictive power (AUC = 0.93) in the validation cohort (n=56) than the panel selected by Byeon et al. (2024)(AUC 0.82). stably represents a principled, error-controlled method for biomarker panel discovery for translation into second cohorts. Availability and implementation stably is available on GitHub (https://github.com/byrnedaniel5-eng/stably); the version used in this study is archived on PyPI (https://pypi.org/project/stably/).

Read PDF

Similar papers

Open access Aug 2026

MSlineaR: An R Package Assessing Linear Behavior to Improve Quality Assurance and Statistical Robustness in Untargeted Metabolomics.

Mass spectrometry-based untargeted metabolomics analyzes complex biological matrices containing thousands of individual features. Linearity is a key analytical parameter in quantitative mass spectrometry, reflecting proportionality between signal intensity and analyte concentration within a defined range. In untargeted...

J. Wiebach, Á. Fernández-Ochoa, Ulrike Bruning et al. · 0 citations
Open access Sep 2026

Dimensionality reduction and metabolite panel derivation in urinary metabolomics based on Random Forest with Gini index feature selection.

These findings highlight the advantage of a unified, supervised tree-based strategy capable of delivering stable classification and interpretable dimensionality reduction across independent untargeted metabolomics platforms, providing a structured and transferable framework for metabolomics-driven biomarker discovery a...

Markus Zetes, Vlad Moisoiu, Carmen Socaciu et al. · 0 citations
Open access Aug 2026

Plasma titration provides a physical ruler for cross-platform proteomics

Plasma proteomics is expanding across platforms and cohorts, and integrating these data for AI demands comparability at the protein level, not merely concordant associations1,2. Affinity and MS platforms use distinct probes (antibodies, aptamers, or peptides) and signal readouts, yielding contradictory cross-platform r...

Yaqing Liu, Haiyan Wang, Yutong Zhang et al. · 0 citations
Open access Aug 2026

Contrastive alignment transfers proteomic predictive signals to metabolomics data

AugMent, a transfer learning framework that uses contrastive learning to encode proteome information into metabolomic representations, supports that contrastive cross-modal learning alignment can be used for transferring disease-relevant information from deeply characterized molecular datasets to substantially larger c...

D. Hu, C. Rohrer, M. Pielies Avelli et al. · 0 citations
#protein folding Open access Sep 2026

Benchmarking of plasma proteomics workflows reveals complementarity of deep MS and affinity-based assays.

This first study to include Olink Reveal in a systematic cross-platform comparison and to directly compare the Orbitrap Astral with the previous-generation Orbitrap Exploris 480 under identical conditions provides a concrete, evidence-based framework for workflow selection and support the co-deployment of deep MS and h...

Jean-Marc Monneuse, Hayat Hage, Célie Da Silva et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.