MACROS establishes a scalable foundation for fully automated structure elucidation, and catalyzes accelerated molecular discovery toward autonomous laboratories, and augments chemists via collaboration to deliver sixfold faster, 40% more accurate elucidation.
Abstract
Following the molecular discovery and synthesis revolutions, scalable automated structure elucidation from routine spectroscopic data remains an outstanding challenge. Despite decades of computational efforts, no existing system achieved reliable reasoning over unseen spectra. Here, we propose MACROS, a multi-agent system automating structure elucidation by emulating expert iterative hypothesis-testing. Trained on 100M simulated and 1.6M experimental spectra-molecule pairs, it natively supports arbitrary combinations of routine spectroscopic techniques. It achieves unprecedented zero-shot generalization to diverse real-world samples, correctly identifying synthetic compounds, natural products and metabolites above 500 Da with 1D NMR. Remarkably, MACROS spontaneously recovers textbook spectroscopic correlations from unassigned data and exhibits emergent chemical intuition such as a ring-first parsing preference, learning fundamental chemical principles rather than memorizing database patterns. MACROS augments chemists via collaboration to deliver sixfold faster, 40% more accurate elucidation. MACROS establishes a scalable foundation for fully automated structure elucidation, and catalyzes accelerated molecular discovery toward autonomous laboratories.
The results show that reframing NMR elucidation as an LLM-guided constrained search, rather than a modeling task, yields substantial gains and suggests a path toward multi-step orchestration frameworks that integrate a variety of tools, models, and domain knowledge to assist in automating spectroscopic analysis.
I. Morales, Damon J. Hinz, Marvin Alberts et al.· 0 citations
Most computationally predicted materials are never synthesized because conventional synthesis optimization is slow, expertise-dependent, and iterative. Here we present a closed-loop framework that automates this expert workflow by placing human tacit knowledge in the loop through a large language model (LLM) that distills synthesis knowledge from the literature, high-throughput hyperspectral imaging for rapid film evaluation, and multi-objective Bayesian optimization guided by experimental feedback. In a paired optimization campaign, LLM-assisted initialization produced more Pareto-optimal samples and higher hypervolume than a Latin hypercube sampling baseline at matched trial counts, and this advantage persisted throughout iterative optimization. We demonstrate the framework by synthesizing the previously unreported perovskite-inspired compound Rb3BiI6 as thin films and validating the optimized films by optical bandgap analysis and X-ray diffraction. The framework transforms synthesis prediction from single-shot recommendation to iterative learning, providing a generalizable strategy to accelerate automated and fully autonomous experimental materials discovery.
Fang Sheng, Steven B. Torrisi, Amanda A. Volk et al.· 0 citations
MolDeTr addresses the spectrum-conditioned inverse problem and extracts spin-system parameters directly from measured 1D 1H NMR spectra, thereby substantially improving chemical-shift prediction precision by one to 2 orders of magnitude compared to existing structure-conditioned approaches.
N. Schmid, Marc Wanner, G. Fischetti et al.· Analytical Chemistry· 1 citation
Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules. We address this challenge by formulating automated structure elucidation as a scalable hypothesis-refinement paradigm that tightly integrates spectral evidence with large-scale molecular priors. To supply structure-resolving NMR signals for multimodal learning, we construct \textbf{QM9SPIN}, a DFT-derived dataset comprising diverse 1D and 2D spectra, including J-coupling, DEPT experiments, and explicit spin--spin interactions. On this foundation, we introduce \textbf{SpectroMol}, a spectrum-to-structure model that proposes chemically valid molecular hypotheses conditioned on multimodal spectral inputs. Complementarily, we develop \textbf{MS-Mol2Mol}, a high-resolution mass-constrained molecular generator that integrates molecular formula, exact mass, and degree of unsaturation within a conditional generative prior trained on 400 million molecules, ensuring global compositional consistency and chemically realistic refinement. The integrated system achieves 93.8\% top-1 accuracy on the simulated benchmark, adapts effectively from simulated to experimental spectra with limited experimental fine-tuning, and further improves experimental predictions through mass-guided refinement, establishing a scalable route toward automated, data-driven organic structure elucidation.
Chengchun Liu, Zhiyuan Yan, Li Yuan et al.· 0 citations
Automation is transforming scientific discovery by enabling systematic exploration of complex hypotheses. Large language models (LLMs) perform well across diverse tasks and promise to accelerate research, but often struggle with logical structures. Here, we present a framework for biological discovery integrating LLM-based agents with laboratory automation, guided by logical scaffolds incorporating symbolic relational learning, structured vocabularies and experimental constraints. This integration improves coherence and reliability in automated workflows. We couple this AI-driven approach to automated cell-culture and metabolomics platforms, enabling integrated hypothesis validation and refinement, yielding a flexible discovery system. The system identified novel interactions in Saccharomyces cerevisiae, including glutamate-induced growth inhibition in spermine-treated cells and aminoadipate's partial rescue of formic-acid stress. All hypotheses, experiments and data are captured in a graph database employing controlled vocabularies. Existing ontologies are extended, and a novel representation of scientific hypotheses is presented using description logics. This work demonstrates the potential for a reliable machine-driven discovery process in systems biology.
Daniel Brunnsåker, Alexander H. Gower, Prajakta Naval et al.· Journal of the Royal Society...· 2 citations
A task-adaptive large reasoning model that integrates chemical knowledge through a synergistic multispecialist architecture, chain-of-thought supervision, and molecule-informed reinforcement learning is presented, demonstrating a versatile multitask framework for knowledge-guided molecular reasoning and design.
Pengfei Liu, Shuang Ge, Xiaobo Wang et al.· Journal of Physical Chemistr...· 0 citations