Skip to content
Open access

makeshift: a lightweight software for accessing and analyzing NMR data and protein dynamics

Aug 2026 · bioRxiv · 0 citations · 24 references
Biology

Abstract

Nuclear magnetic resonance (NMR) spectroscopy yields rich residue-level information on biomolecular dynamics and chemical environments, two frontiers for quantitative predictive methods in biochemistry. Decades of data are publicly archived in the Biological Magnetic Resonance Data Bank (BMRB)1, yet in practice, this information remains difficult to access and interpret at scale and within computational workflows. Here we present makeshift, an open-source Python package for accessing, curating, and analyzing NMR datasets. Users can readily retrieve and parse BMRB entries and perform essential analyses such as chemical shift re-referencing, secondary structure propensity prediction, and interpretation of relaxation datasets for dynamics. We re-implemented several widely-used NMR data calculations which were not open-source or available in Python and validated our implementations against the original implementations. By integrating data access, processing, and analysis into a single Python interface, makeshift lowers the barrier for reproducible, scalable analysis and machine learning applications using biomolecular NMR data.

Read PDF

Similar papers

Open access Aug 2026

Inferring protein ensembles directly from NOESY spectra

A quantitative scoring framework for comparing experimental and back-calculated observables is introduced and combined with regularized ensemble selection and Monte Carlo simulated annealing to provide direct inference of protein ensembles within a flexible ensemble-selection architecture incorporating multiple classes of NMR observables.

Murray Coles · 0 citations
Open access Aug 2026

Learning millisecond protein dynamics from what is missing in NMR spectra.

Many proteins' biological functions rely on interconversions between multiple conformations occurring at micro- to millisecond (µs-ms) timescales. A lack of standardized, large-scale experimental data has hindered obtaining a more predictive understanding of these motions. After curating >100 Nuclear Magnetic Resonance (NMR) relaxation datasets, we realized an observable for µs-ms dynamics might be hiding in plain sight. Millisecond dynamics can cause NMR signals to broaden beyond detection, leaving some residues not assigned in the chemical shift datasets of ~10,000 proteins deposited in the Biological Magnetic Resonance Data Bank (BMRB) 1. We made the bold assumption that residues missing assignments are exchange-broadened due to µs-ms motions and trained various deep learning models to predict missing assignments. Strikingly, these models also predict exchange measured via NMR relaxation experiments, indicative of µs-ms dynamics. The best of these models, which we named Dyna-1, leverages an intermediate layer of the multimodal language model ESM-32. Notably, dynamics directly linked to biological function, including enzyme catalysis and ligand binding, are particularly well predicted by Dyna-1, which parallels our findings that residues experiencing µs-ms exchange are more conserved. We anticipate the datasets and models presented here will be transformative in unlocking the common language of dynamics and function.

Hannah K. Wayment-Steele, Gina El Nesr, Ramith Hettiarachchi et al. · 0 citations
Open access Aug 2026

SPINDLE: Unlocking protein dynamics from single-field NMR relaxation data using a deep learning ensemble

SPINDLE, an ensemble of deep neural networks trained on a large synthetic set of NMR relaxation data predicts both fast and slow timescale dynamics parameters from a single set of three relaxation datasets using the ensemble for error estimation, and demonstrates a strong correlation to ground truth dynamics parameters on synthetic benchmarks.

Olivia E. Krise, Michael P. Latham · 0 citations
Open access Sep 2026

Accessible hybrid DFT-quality NMR crystallography via gas-phase machine learning interatomic potentials

Nuclear magnetic resonance (NMR) crystallography is a robust method for structure determination, but its reliance on density functional theory (DFT) calculations for geometry refinement limits its speed and accessibility. Recent machine-learning predictors such as ShiftML3 can evaluate magnetic shieldings in seconds, but still need high-quality crystal geometries that are normally obtained from slow DFT optimisations. Here, we demonstrate that machine learning interatomic potentials (MLIPs) can eliminate this bottleneck for organic crystals. Interestingly, models trained on gas-phase molecules at the hybrid-DFT ωB97M-V level (OMol25 dataset)—such as UMA-omol and MACE-Polar-1—deliver structural quality rivaling hybrid periodic-DFT, without requiring large computational resources. The resulting MLIP/ShiftML3 workflow reduces computational costs by at least three orders of magnitude, making high-accuracy NMR crystallography accessible without HPC infrastructure, and opening the door to fast molecular dynamics simulations at the hybrid DFT level. Combined with sensitivity enhancement via dynamic nuclear polarization (DNP), we demonstrate the approach on l-histidine, extracting 13C and 15N chemical shift tensors and using 1H chemical shifts at natural isotopic abundance to discriminate between its monoclinic and orthorhombic polymorphs.

Shubha S. Gunaga, Rob Schurko, Sean T. Holmes et al. · 2 citations
Aug 2026

Principal Component Analysis Based Deconvolution of NMR Chemical Shifts for Objective Biomolecular Interaction Mapping

Identifying ligand-binding interfaces from NMR chemical shift perturbation data is a cornerstone of structural biology, yet it remains hampered by subjective thresholding and the loss of directional information in conventional scalar metrics. Traditional methods rely on empirical weighting factors that lack physical universality and ignore critical parameters like exchange-induced line broadening. Here, we introduce PALI (principal component analysis for ligand interactions), an objective and PCA-based framework designed to standardize multivariate NMR analysis. By implementing Z-score standardization, PALI replaces arbitrary constants with a rigorously data-driven statistical framework, preserving the multidimensionality of spectral changes. We demonstrate that PALI effectively filters stochastic noise and identifies binding hotspots, including features such as intermediate exchange that are often obscured in conventional 1D plots. Validation across diverse systems, from structured proteins to complex dynamic assemblies and intrinsically disordered regions/proteins, proves that PALI provides a robust, reproducible, and automated solution for interaction mapping. PALI is freely available as a Web-based dashboard, bridging the gap between advanced multivariate statistics and routine structural biology workflows.

Min June Yang, Joonhyeok Choi, Hyeonjun Lee et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.