Skip to content

Prediction of mass spectra using large chemical language models and verification of adaptability in data-scarce domains

Aug 2026 · Applied Physics Letters · 0 citations · 12 references

TL;DR

Results demonstrate that MolFormer-XL, which combines pre-trained molecular representations with a Transformer-based architecture and learned SMILES embeddings, provides a promising approach for transfer under severe domain-specific data scarcity in environmental mass spectrometry.

Abstract

Machine learning methods for predicting the electron ionization mass spectra from molecular structures have shown promise for environmental chemical identification, but their performance under domain-specific data scarcity remains poorly understood. We systematically compare a conventional multilayer perceptron model (NEIMS) with a Transformer-based chemical foundation model (MolFormer-XL) for the electron ionization mass spectrometry spectrum prediction under controlled few-shot conditions. Using fluorine-containing molecules as a broader proxy domain, including a PFAS-like subset, motivated by the practical challenge of detecting novel fluorinated contaminants with limited reference data, we vary the number of domain-specific training examples from 5 to 175 while maintaining fixed validation and test sets. Across all few-shot conditions and three of four evaluation metrics (weighted cosine similarity, intensity-weighted precision, and top-10 precision), MolFormer-XL consistently outperforms NEIMS, while intensity-weighted recall remains comparable between the two models. The largest performance gaps are observed in extreme data-scarcity regimes. These results demonstrate that MolFormer-XL, which combines pre-trained molecular representations with a Transformer-based architecture and learned SMILES embeddings, provides a promising approach for transfer under severe domain-specific data scarcity in environmental mass spectrometry.

View source

Similar papers

Open access Sep 2026

Learning from tandem mass spectra at scale with a self-supervised foundation model for proteomics

Mass spectrometry-based proteomics increasingly relies on machine learning, yet existing models are trained for defined supervised tasks such as peptide identification, de novo sequencing or fragment intensity prediction, limiting transfer across datasets, instruments and acquisition methods. Here we present InstaNovo-...

M. Nieuwoudt, Marco Reverenna, Divanisha Patel et al. · 0 citations
Open access Aug 2026

In-Context Learning Meets Small Molecule Property Prediction: Benchmarking Novel Machine Learning Approaches.

Recently, a new category of machine learning approaches for tabular data has emerged: tabular foundation models (TFM), based on in-context learning. A TFM is a neural network (usually a transformer) pretrained primarily on synthetic data. Its input is an entire data set: features and labels for training records, along...

D. Matyushin, A. Sholokhova · 0 citations
Aug 2026

FullMS2Former: A Transformer Model for Near-Complete Prediction of Peptide Tandem Mass Spectra.

Accurate prediction of peptide tandem mass spectra is essential for confident peptide identification. Existing deep learning models achieve high similarity to experimental spectra but remain constrained by fixed fragment dictionaries of common fragment ions, limiting their ability to capture uncommon or previously unre...

Justin R. Zhang, Zhongqi Zhang · 0 citations
Preprint Sep 2026

Simulation-Supervised Foundation Models for Retention Time Prediction in High-Performance Liquid Chromatography beyond Experimental Data Coverage

Accurate prediction of high-performance liquid chromatography (HPLC) retention times (RTs) across diverse molecules and chromatographic methods remains challenging because experimental training data cover only a limited region of chemical and method spaces. Here, we develop FUSE-RT (Foundation model Unifying Simulation...

Stephen Wu, Yuxuan Han, Yasuhiro Mito et al. · 0 citations
Open access Jul 2026

Foundation Models for Liquid Chromatography–High-Resolution Mass Spectrometry: A New Era beyond Labeled Datasets

Liquid chromatography coupled to high-resolution mass spectrometry (LC–HRMS) is a widely used analytical technique for characterizing the chemical composition of organic samples. Due to its high sensitivity and ability to detect thousands of chemical features in a single run, untargeted LC–HRMS experiments generate hig...

A. Carnoli, F. Padilla-González, Leonieke M. van den Bulk et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.