From metabolomics dark matter to bioactivity: a computational framework for prioritizing food-derived bioactive compounds
Despite the growing scale of untargeted food metabolomics, scalable approaches for translating raw LC-MS profiles into functionally prioritized, experimentally testable candidates remain limited. Most detected features remain structurally unresolved, and the few that are annotated cannot be reliably connected to whole-food biological activity through additive models alone. Here we present a multi-layer prioritization framework that addresses this bottleneck across 322 commonly consumed foods in the United States, comprising 30,343 small molecules. A retention-index model ( R ² = 0.95, PCC = 0.97) expanded putative structural coverage from 7.4% to over 78% of detected features, assigning confidence-tiered structural proxies through formula-constrained matching against 11.8 million PubChem compounds. A two-stage large language model workflow harmonized 39,158 assay protocols spanning 119,633 chemical-assay measurements from ChEMBL into standardized bioactivity classes, enabling compound-level evidence to be propagated across the food composition space. The framework was applied to antioxidant activity as a primary, well-characterized endpoint while additional bioactivity domains (anti-inflammatory, anticancer, antidiabetic, antiviral, and antibacterial effects) were incorporated as compound-level evidence layers to support metabolite prioritization. Integrating concentration-aware activity mapping with a machine learning model trained on whole-food antioxidant measurements yielded strong predictive performance (R² = 0.73, PCC = 0.85) and prioritized metabolites most strongly associated with model-predicted food-level antioxidant capacity via SHAP attribution. Structure-based analysis using Boltz-2 against the antioxidant sensor provided a final mechanistic triage layer, distinguishing high-priority candidates from lower-ranked compounds. Applied across 322 foods, this funnel reduces tens of thousands of untargeted LC-MS features to a confidence-ranked shortlist for targeted wet-lab validation and provides a generalizable strategy for translating food metabolomics into biologically interpretable and actionable hypotheses. This framework establishes a scalable foundation for transforming food composition data into actionable biological insight, advancing the integration of food, health, and data science. More broadly, these results highlight limitations of traditional nutrient-centric frameworks and support a shift toward modeling molecular patterns and compositional signatures to better capture the functional properties of foods.