Aug 2026· Metabolomics· Vol 22· 0 citations· 48 references
Medicine
TL;DR
A two-step analytical approach was developed to systematically prioritize and interpret unannotated metabolites using plasma LC–MS/MS data from pregnant women with obesity as a biologically relevant test dataset, providing a scalable framework for prioritizing dark matter metabolites in untargeted metabolomics.
Abstract
Untargeted metabolomics often results in a significant portion of unannotated metabolites, or “metabolic dark matter,” which hinders biological interpretation. A two-step analytical approach was developed to systematically prioritize and interpret unannotated metabolites using plasma LC–MS/MS data from pregnant women with obesity as a biologically relevant test dataset. The first step involved clustering 1,021 known metabolites into ten structurally coherent groups based on the Tanimoto similarity, thus defining the biologically relevant chemical space of the dataset. These metabolites were further characterized by Absorption, Distribution, Metabolism, and Excretion (ADME) profiling, protein target prediction, molecular docking and Kyoto Encyclopedia of Genes and Genomes pathway mapping analysis, to establish biological plausibility and functional perspective. Candidate structures for 1,836 unannotated features were retrieved from PubChem using molecular formula and molecular weight matching within a ±0.5 Da tolerance. This search yielded 569,115 candidate structures, of which 368,197 unique structures were retained after curation. Tanimoto coefficient filtering reduced the candidate pool to 19,868 structurally plausible candidates, and retention time-based prioritization further refined this set to 418 high confidence candidate annotations, including 83 database-supported candidates identified through HMDB and LIPID MAPS structure database cross-referencing. RT-based prioritization effectively distinguished positional isomers sharing the same molecular formula by incorporating agreement between predicted and experimentally observed retention times. This improved discrimination among structurally similar candidates, expanded metabolite annotation confidence, and provided a scalable framework for prioritizing dark matter metabolites in untargeted metabolomics. Clustered-based workflow integrating chemical similarity and retention time to prioritize and annotate unknown metabolites
Despite the growing scale of untargeted food metabolomics, scalable approaches for translating raw LC-MS profiles into functionally prioritized, experimentally testable candidates remain limited. Most detected features remain structurally unresolved, and the few that are annotated cannot be reliably connected to whole-food biological activity through additive models alone. Here we present a multi-layer prioritization framework that addresses this bottleneck across 322 commonly consumed foods in the United States, comprising 30,343 small molecules. A retention-index model (
R
² = 0.95, PCC = 0.97) expanded putative structural coverage from 7.4% to over 78% of detected features, assigning confidence-tiered structural proxies through formula-constrained matching against 11.8 million PubChem compounds. A two-stage large language model workflow harmonized 39,158 assay protocols spanning 119,633 chemical-assay measurements from ChEMBL into standardized bioactivity classes, enabling compound-level evidence to be propagated across the food composition space. The framework was applied to antioxidant activity as a primary, well-characterized endpoint while additional bioactivity domains (anti-inflammatory, anticancer, antidiabetic, antiviral, and antibacterial effects) were incorporated as compound-level evidence layers to support metabolite prioritization. Integrating concentration-aware activity mapping with a machine learning model trained on whole-food antioxidant measurements yielded strong predictive performance (R² = 0.73, PCC = 0.85) and prioritized metabolites most strongly associated with model-predicted food-level antioxidant capacity via SHAP attribution. Structure-based analysis using Boltz-2 against the antioxidant sensor provided a final mechanistic triage layer, distinguishing high-priority candidates from lower-ranked compounds. Applied across 322 foods, this funnel reduces tens of thousands of untargeted LC-MS features to a confidence-ranked shortlist for targeted wet-lab validation and provides a generalizable strategy for translating food metabolomics into biologically interpretable and actionable hypotheses. This framework establishes a scalable foundation for transforming food composition data into actionable biological insight, advancing the integration of food, health, and data science. More broadly, these results highlight limitations of traditional nutrient-centric frameworks and support a shift toward modeling molecular patterns and compositional signatures to better capture the functional properties of foods.
Pranav Gupta, Michael Gunning, Selena Ahmed et al.· npj Science of Food· 0 citations
The convergence of docking, dynamics, and free-energy results prioritized PM2, PM3, and PM4 as promising MAP3K8 hit candidates, which require further experimental validation and lead optimization.
M. Islam, A. Iqbal, M. A. Ali et al.· Molecular diversity· 0 citations
Natural products bearing novel skeletons expand accessible chemical space and may inspire future biological discovery, yet their identification still relies largely on chance. This limitation stems largely from prevailing strategies for representing and interpreting MS2 data and not simply from spectral quality or algorithmic performance. Current workflows treat MS2 spectra as pairwise comparable entities, with structural relationships inferred from similarity, favoring analogues of known scaffolds while obscuring global structural organization. Here, we propose a feature-structured analytical paradigm that treats MS2 data as an integrated signal system shaped by molecular skeletons. By reorganizing spectra into fragment-feature representations and applying non-negative matrix factorization, structural information becomes directly observable at the data set level, revealing a skeleton-level chemical space (SLECS) in which compounds organize according to underlying skeletal features rather than pairwise similarity. Evaluation across diverse data sets shows that the SLECS-based framework resolves skeleton-level organization beyond conventional similarity-based approaches and enables systematic identification of structurally distinct regions. Application to Pilea cavaleriei led to the targeted isolation of seven sesquiterpenes (1-7), including compounds 1 and 2, featuring an unprecedented bicyclo[6.3.1] skeleton, thereby enabling targeted discovery of novel skeletons and expanding the pool of unexplored scaffolds for future biological evaluation. This work establishes a new representation paradigm for MS2 data analysis that complements similarity-based approaches and offers a scalable strategy for structure-oriented exploration of chemical space and the discovery of novel natural product scaffolds.
Licorice (Glycyrrhiza) is a medicinal plant widely used in approximately 70% of traditional Japanese Kampo formulations and is known to produce a wide array of specialized metabolites with diverse pharmacological properties. Although hundreds of metabolites have been reported, the overall chemical diversity of Glycyrrhiza remains poorly characterized. Here, using mass spectrometry data obtained from fully 13C-labeled leaves and roots of Glycyrrhiza uralensis and Glycyrrhiza glabra, we determined the carbon number, followed by the molecular formula and substructure prediction in combination with MS/MS similarity-based molecular networking. After excluding redundant ions, including isotopic peaks, adducts, and in-source fragments, we extracted 3060 unique metabolite features with assigned carbon numbers. Among these, substructure information was assigned to 1015 features (33%) across the four plant tissues, revealing tissue-specific metabolome profiles. Furthermore, we discovered five previously unreported homopipecolic acid-conjugated flavonoids in the roots of G. uralensis and G. glabra and Glycine max, another member of the Fabaceae family. Two of these compounds were structurally characterized using nuclear magnetic resonance spectroscopy. We further proposed a biosynthetic route involving a spontaneous reaction between 1-piperideine and malonyl glycoside substrates and confirmed the formation of the conjugated product using authentic standards.
Keita Sawai, Yoshimasa Todoroki, Shohei Nakamukai et al.· Journal of Natural Products· 0 citations
Introduction Streptomyces species represent an important source of bioactive natural products, yet systematic genome-guided prioritization of metabolites targeting cyclooxygenase-2 (COX-2/PTGS2) remains limited. This study aimed to investigate the biosynthetic potential of Streptomyces sp. VITGV156 (MCC 4965) using an integrated genome mining and computational drug discovery pipeline. Methods Whole-genome sequencing, functional annotation, antiSMASH v7.0.1-based biosynthetic gene cluster (BGC) prediction, LC-MS/MS metabolomic profiling, SwissADME analysis, target prediction, disease association mapping, molecular docking against PTGS2 (PDB: 5IKR), and PASS bioactivity prediction were performed to prioritize putative bioactive metabolites. Results Genome analysis identified 29 predicted biosynthetic gene clusters, including clusters associated with geosmin, ectoine, albaflavenone, hopene, coelichelin, and SapB, together with several cryptic clusters exhibiting low similarity to known pathways. LC-MS/MS metabolomic profiling provided experimental support for active secondary metabolite production under the cultivation conditions employed. Computational prioritization identified PTGS2 (COX-2) as a biologically relevant target. Molecular docking demonstrated favorable binding affinities and interaction profiles for several predicted metabolites within the PTGS2 catalytic pocket. PASS analysis further suggested potential anticancer-related biological activities that require experimental validation. Discussion These findings demonstrate the utility of integrating genome mining, metabolomic profiling, and computational drug discovery for prioritizing natural-product candidates. Streptomyces sp. VITGV156 (MCC 4965) represents a promising source of biosynthetic diversity and provides a genome-guided framework for identifying putative COX-2-targeting natural products for future experimental validation rather than confirming metabolite production or biological activity.
Veilumuthu Pattapulavar, Saranyadevi Subburaj, Sathiyabama Ramanujam et al.· Frontiers in Bioinformatics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.