Skip to content
#protein folding Open access

Enhanced classification and identification of bacterial and viral microorganisms by integration of MALDI-TOF mass spectrometry with artificial intelligence

Aug 2026 · Scientific Reports · Vol 16 · 0 citations · 53 references
Medicine

TL;DR

The Extra Trees Classifier consistently achieved the highest average accuracy and F1-score in both Gram type classification and species-level identification, demonstrating superior generalization across datasets.

Abstract

The accurate and rapid identification of bacterial pathogens is essential in clinical setups and medical biodefense. Matrix-Assisted Laser Desorption Ionization Time-Of-Flight (MALDI-TOF) mass spectrometry has emerged as a powerful tool for fast and reliable microbial identification. This study assesses the performance of eight Machine Learning (ML) and two Deep Learning (DL) models trained using 5-fold cross validation in classifying microorganisms in a series of experiments based on MALDI-TOF mass spectra (n = 255) generated internally from seven cultured bacteria and five viral agents. The results showed that up to seven models achieved consistently robust classification across various tasks, including binary classification of viruses from bacteria and Gram-positive from Gram-negative bacteria, as well as multi-class classification of individual bacterial and viral species. Subsequently, we cross-validated three from our top-performing models, Extra Trees Classifier, Support Vector Classifier and 1-D Convolutional Neural Network, on Gram type and multi-class classification of individual bacterial species against an external dataset selected from a large MALDI-TOF MS database of highly pathogenic bacteria (curated by Robert Koch Institute). Among these, the Extra Trees Classifier consistently achieved the highest average accuracy and F1-score in both Gram type classification and species-level identification, demonstrating superior generalization across datasets. Its ensemble architecture proved particularly effective in capturing subtle spectral patterns associated with microbial cell wall composition and protein expression profiles. These results underscore the strong potential of this model to enhance MALDI-TOF-based classification frameworks in clinical microbiology and biodefense applications.

Read PDF

Similar papers

Open access Aug 2026

Integrating multi-species and multi-antibiotic resistance classification with MALDI-TOF: a deep learning approach to predict AMR

Antimicrobial resistance (AMR) poses a significant global health threat, impacting clinical treatments, agriculture, and public health. Although mass spectrometry techniques like MALDI-TOF provide opportunity for rapid AMR detection, current state-of-the-art models, such as MSDeepAMR, are limited to single-label classification, requiring separate models for each bacterium-antibiotic combination. These approaches struggle with challenges such as class imbalance, incomplete labels, and poor generalization across bacterial strains and antibiotics. This study addresses these limitations by introducing a novel multi-label, multi-bacteria classification framework (MLMBC) that simultaneously predicts AMR across multiple bacterial species and antibiotics using MALDI-TOF mass spectrometry data and Convolutional Neural Networks (CNNs), hereby establishing a solid foundation for the development of robust and rapid diagnostic tools to address the growing threat of multidrug-resistant bacteria. Utilizing CNNs in combination with transfer learning and semi-supervised classification techniques, such as self-training, pseudo-labeling and kNN label propagation, our approach improves accuracy and overcomes dataset limitations. Results demonstrate that the proposed approach with pseudo-labeling consistently outperforms single- and multi-label baselines; for example, it achieves AUROC \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\ge$$\end{document} 0.94 for E. coli, K.pneumoniae and S.aureus - Ceftriaxone pairs. Statistical tests support the effectiveness of the proposed approach, demonstrating that simultaneous multi label, multi-bacteria modeling constitutes a reliable and scalable strategy for antibiotic resistance prediction in clinically relevant settings.

Leila Aro-Sati, Xaviera A. López-Cortés, J. M. Troncoso et al. · 0 citations
Open access Jul 2026

Machine Learning-Based Prediction of Antimicrobial Resistance in Escherichia coli from MALDI-TOF Mass Spectrometry Data

Objectives: To assess the feasibility and reproducibility of predicting antimicrobial resistance (AMR) in Escherichia coli from MALDI-TOF mass spectrometry data using a standardized, open-source machine learning (ML) workflow, we systematically compared four ML algorithms, evaluated the impact of culture conditions, extract storage, and spectral preprocessing on model performance, and validated results through nested cross-validation with statistical significance testing. Methods: A total of 282 clinical E. coli isolates were analyzed. Two MALDI-TOF MS datasets were generated from freshly cultured extracts (T1) and recultured isolates one year later (T3), yielding 4468 spectra. A third dataset from the T1 extracts stored at −20 °C for one year (T2) was evaluated for spectral stability but excluded from primary modeling likely due to storage-induced degradation. Protein spectra (m/z 2000–15,000) were preprocessed using an in-house developed MALDI-TOF preprocessing pipeline (MTPP) comprising variance stabilization, Savitzky–Golay smoothing, SNIP baseline correction, TIC normalization, LOWESS alignment, and MAD-based peak detection (SNR ≥ 3), yielding 121 m/z features. Four classifiers—Random Forest (RF), Logistic Regression, Support Vector Machine, and Gradient Boosting—were trained to predict resistance to 11 antibiotics using nested cross-validation: outer GroupShuffleSplit (5-fold, isolate-level) for evaluation and inner GroupKFold for recursive feature elimination (RFECV) and hyperparameter tuning (RandomizedSearchCV). Classification thresholds were optimized via the precision–recall curve. Model performance was assessed using AUROC, AUPRC, F1-score, Matthews Correlation Coefficient (MCC), and bootstrap 95% confidence intervals (1000 replicates). Pairwise model comparisons were tested with McNemar’s chi-squared test. Results: Among the 12 antibiotics included in the analysis (meropenem excluded for absence of resistance), resistance prevalence ranged from 1.1% (colistin) to 59.9% (amoxicillin). Colistin was subsequently also excluded from ML modeling due to insufficient resistant isolates (n = 3), leaving 11 antibiotics for prediction. The best predictive performance was observed for ciprofloxacin (AUROC 0.76 [95% CI 0.74–0.77]; F1 0.54; MCC 0.38) and ceftazidime (AUROC 0.68 [0.65–0.71]; F1 0.36; MCC 0.29), using 13 and 37 RFECV-selected features, respectively. Amoxicillin achieved the highest F1-score (0.76), driven by high recall (0.98) but modest AUROC (0.58). No meaningful predictive signal was detected for amikacin, cefepime, or tigecycline (AUROC ≤ 0.57, F1 ≤ 0.17), attributable to extreme class imbalance, and no robust multi-peak resistance signature was detected in this dataset. McNemar’s test confirmed that RF significantly outperformed Logistic Regression for all antibiotics (p < 0.01), while Gradient Boosting performed comparably to RF for ciprofloxacin (p = 0.17) and ceftazidime (p = 0.28). Frozen extracts (T2) produced lower spectral similarity and were excluded from model training; the aligned T1+3 dataset yielded the most stable performance across metrics. Conclusions: Machine learning analysis of MALDI-TOF spectra enables reproducible AMR prediction for selected antibiotics in E. coli, with ciprofloxacin and ceftazidime showing the strongest signal. Nested isolate-level cross-validation, multi-model comparison with statistical testing, and open-source code provide a transparent, reproducible foundation for integrating ML-assisted MALDI-TOF analysis into diagnostic AMR surveillance. Extract storage at −20 °C degrades spectral quality and should be avoided in ML training workflows.

N. Versmessen, Marieke Mispelaere, R. Vanstokstraeten et al. · 0 citations
Open access Aug 2026

Portable Raman spectroscopy combined with machine learning for rapid and label-free recognition of five pathogenic bacteria

Introduction Antimicrobial resistance (AMR) continues to rise globally, highlighting the need for rapid, label-free, and cost-effective bacterial identification methods. In this proof-of-concept study, portable Raman spectroscopy combined with machine learning was used to identify five clinically relevant bacterial species: Escherichia coli, Pseudomonas aeruginosa, Staphylococcus aureus, Porphyromonas gingivalis, and Streptococcus mutans. Methods Raman spectra were acquired from cultured, washed, PBS-resuspended, and OD-standardized bacterial suspensions. After SNIP baseline correction, binary and five-class classification models were constructed using Auto-Sklearn with eight algorithms: ADB, ET, GB, LDA, SVM, MLP, PA, and QDA. Model performance was evaluated using accuracy, precision, recall, F1-score, MCC, and ROC-AUC. Results Pairwise binary classification showed variable performance among bacterial pairs. The best result was obtained for P. gingivalis versus S. aureus, with a testing accuracy of 98.3%, precision of 0.984, recall of 0.983, F1-score of 0.983, MCC of 0.967, and ROC-AUC of 1.000. Five-class classification was more limited, with LDA achieving the highest testing accuracy of 60.1%, MCC of 0.506, and ROC-AUC of 0.867. Discussion These findings support the feasibility of portable Raman spectroscopy combined with machine learning for bacterial recognition under standardized sample conditions.

Shisheng Cao, Ran Pang, Yongqiang Chen et al. · 0 citations
Open access Jul 2026

Machine learning–assisted surface-enhanced Raman spectroscopy for multiple and rapid screening of foodborne pathogenic bacteria

To address the challenge of mixed contamination of foodborne pathogenic bacteria in food, in this study, a machine learning (ML) assisted surface-enhanced Raman scattering (SERS) sensing platform was proposed for multiple and rapid screening of foodborne pathogenic bacteria. 4-Mercaptophenylboronic acid-functionalized gold nanoparticles (AuNPs@4-MPBA) was introduced as SERS substrate and the molecular recognition and electromagnetic enhancement mechanisms for target analytes were elucidated by theoretical simulation. The sensing platform achieved efficient identification and differentiation of single and mixed contaminations of Escherichia coli, Salmonella typhimurium, Shigella, Listeria monocytogenes, and Staphylococcus aureus in three tea samples. The ability of the model to generalize across varying conditions was evaluated using a mixed spectral dataset from different tea sample matrices. After spectral preprocessing, the Random Forest (RF) model was used to classify 31 samples, achieving a classification accuracy of 97.34% and an out-of-bag accuracy of 95.98%, demonstrating excellent stability and generalization capability. The proposed approach enabled the rapid and accurate classification of foodborne pathogenic bacteria in complex food matrices, offering a promising strategy for rapid food safety emergency screening.

Simin Dai, Ceping Yin, Xuejing Fan et al. · 0 citations
Open access Aug 2026

An Interpretable Multi-Objective Machine Learning Framework for In Silico Prioritization of Anti-Staphylococcus aureus Antimicrobial Peptides

Background: Staphylococcus aureus, including methicillin-resistant lineages, is a leading cause of device- and catheter-related infection, and rising resistance motivates the search for antimicrobial peptides (AMPs) with strong anti-staphylococcal activity and low host toxicity. Machine learning can prioritize candidate peptides. However, the literature-derived AMP datasets are prone to homology-driven optimism, and computational studies frequently overstate their translational reach. Methods: We curated 4007 deduplicated S. aureus-active AMP records and 582 binary-labeled hemolysis records. Each peptide was encoded with a transparent 538-dimensional physicochemical and compositional feature vector. Five classifiers and five regressors were evaluated for four endpoints (potency classification, log10 MIC regression, hemolysis classification, normalized hemolytic index) under both conventional random 5-fold cross-validation and homology-aware cross-validation, in which sequences were clustered by 3-mer similarity and whole clusters were confined to single folds. Class imbalance was handled by class weighting. Model behavior was interpreted with SHAP and Fisher-exact k-mer enrichment, and candidates were ranked by a multi-objective score that combines the independently trained heads. Results: Under homology-aware validation, performance was lower than under random splitting, as expected. Potency classification reached an AUROC of about 0.71 (Random Forest), compared with 0.797 under random cross-validation. Hemolysis classification remained strong at AUROC 0.90 (95% CI 0.88 to 0.93), which indicates that its high accuracy is not a homology leakage artifact. MIC regression was modest (homology-aware R2 0.17, Spearman ρ 0.39) and is therefore treated only as a rank-ordering signal. SHAP and k-mer analyses recovered interpretable structure–activity relationships. Net positive charge and amphipathicity drove potency, whereas bulk hydrophobicity drove hemolysis. Applying the pipeline to a generated pool prioritized 20 candidates. Nearest-neighbor analysis shows that these are close optimized variants of known potent scaffolds, with a median identity of 95% to a known peptide, rather than novel sequences. Conclusions: We present an interpretable, honestly benchmarked multi-objective pipeline that optimizes known anti-S. aureus AMP scaffolds toward lower predicted hemolysis. The prioritized peptides are computational hypotheses for future synthesis and experimental testing. Their low predicted hemolysis reflects a selection criterion rather than validated safety, and cross-species selectivity was not assessed.

Jianguo Xu, Donghua Yang, Q. Zheng et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Google DeepMind Blog Nov 25, 2025

AlphaFold: Five years of impact

Explore how AlphaFold has accelerated science and fueled a global wave of biological discovery.