Aug 2026· Applied Physics Letters· 0 citations· 12 references
TL;DR
Results demonstrate that MolFormer-XL, which combines pre-trained molecular representations with a Transformer-based architecture and learned SMILES embeddings, provides a promising approach for transfer under severe domain-specific data scarcity in environmental mass spectrometry.
Abstract
Machine learning methods for predicting the electron ionization mass spectra from molecular structures have shown promise for environmental chemical identification, but their performance under domain-specific data scarcity remains poorly understood. We systematically compare a conventional multilayer perceptron model (NEIMS) with a Transformer-based chemical foundation model (MolFormer-XL) for the electron ionization mass spectrometry spectrum prediction under controlled few-shot conditions. Using fluorine-containing molecules as a broader proxy domain, including a PFAS-like subset, motivated by the practical challenge of detecting novel fluorinated contaminants with limited reference data, we vary the number of domain-specific training examples from 5 to 175 while maintaining fixed validation and test sets. Across all few-shot conditions and three of four evaluation metrics (weighted cosine similarity, intensity-weighted precision, and top-10 precision), MolFormer-XL consistently outperforms NEIMS, while intensity-weighted recall remains comparable between the two models. The largest performance gaps are observed in extreme data-scarcity regimes. These results demonstrate that MolFormer-XL, which combines pre-trained molecular representations with a Transformer-based architecture and learned SMILES embeddings, provides a promising approach for transfer under severe domain-specific data scarcity in environmental mass spectrometry.
Mass spectrometry-based proteomics increasingly relies on machine learning, yet existing models are trained for defined supervised tasks such as peptide identification, de novo sequencing or fragment intensity prediction, limiting transfer across datasets, instruments and acquisition methods. Here we present InstaNovo-...
M. Nieuwoudt, Marco Reverenna, Divanisha Patel et al.· bioRxiv· 0 citations
Recently, a new category of machine learning approaches for tabular data has emerged: tabular foundation models (TFM), based on in-context learning. A TFM is a neural network (usually a transformer) pretrained primarily on synthetic data. Its input is an entire data set: features and labels for training records, along...
D. Matyushin, A. Sholokhova· Journal of Chemical Informat...· 0 citations
Accurate prediction of peptide tandem mass spectra is essential for confident peptide identification. Existing deep learning models achieve high similarity to experimental spectra but remain constrained by fixed fragment dictionaries of common fragment ions, limiting their ability to capture uncommon or previously unre...
Justin R. Zhang, Zhongqi Zhang· Analytical Chemistry· 0 citations
Accurate prediction of high-performance liquid chromatography (HPLC) retention times (RTs) across diverse molecules and chromatographic methods remains challenging because experimental training data cover only a limited region of chemical and method spaces. Here, we develop FUSE-RT (Foundation model Unifying Simulation...
Stephen Wu, Yuxuan Han, Yasuhiro Mito et al.· 0 citations
Liquid chromatography coupled to high-resolution mass spectrometry (LC–HRMS) is a widely used analytical technique for characterizing the chemical composition of organic samples. Due to its high sensitivity and ability to detect thousands of chemical features in a single run, untargeted LC–HRMS experiments generate hig...
A. Carnoli, F. Padilla-González, Leonieke M. van den Bulk et al.· Analytical Chemistry· 0 citations