Skip to content

Active-Learning Discovery of Superionic Compositions Using High-Throughput EIS and Structure-Aware Descriptors

Jul 2026 · ECS Meeting Abstracts · 0 citations

Abstract

We introduce an active-learning framework that closes the loop between high-throughput EIS measurements and structure-aware composition descriptors to discover superionic candidates under realistic processing constraints. Starting from a small seed set, Gaussian-process and tree-based models propose batched experiments that maximize information gain on conductivity and activation energy while enforcing uncertainty-aware Kramers–Kronig quality gates. Descriptor families integrate interpretable features: ionic radius mismatch, framework softness, site connectivity from simple graph-derived motifs, and processing proxies (grain size from Scherrer, porosity, interphase penalty terms). We demonstrate rapid convergence to high-conductivity regions in multi-component chalcogenide and halide spaces using the automated multi-site EIS workflow described separately. Across three material spaces, the approach reduces experiments ~3× versus grid sampling while yielding candidates with improved conductivity at moderate temperatures and stable impedance upon cycling. We release a lightweight, reproducible stack (metadata schema, analysis notebooks, and synthetic datasets) to encourage community benchmarking without proprietary infrastructure. The result is a pragmatic path to self-driving electrolyte discovery that prioritizes experimental tractability and interpretability—features that matter for industrial translation and cross-lab reproducibility. Keywords: active learning; Bayesian optimization; EIS QC; interpretable descriptors; high-throughput screening; solid electrolytes

View source

Similar papers

Jul 2026

Mechanism-Guided Catalyst Discovery for Methane C-H Activation via Structure-Aware Multisource Transfer Learning.

A mechanism-driven approach to alleviate data dependence and develop a multisource transfer learning (MS-TL) framework that leverages the knowledge embedded in abundant adsorption data sets while accurately capturing local structural dependence, enabling a deep fusion of multidimensional thermodynamic knowledge while preserving local structural information.

Wangqiang Lin, Huiyan Zhang, Jinxin Sun et al. · 0 citations
Jul 2026

Data-Driven Exploration of the Polyethylene Catalyst Chemical Space via Machine Learning.

A data-driven framework combining explainable machine learning (ML) with large-scale virtual library generation with large-scale virtual library generation is presented, establishing a practical route from experimental data to actionable catalyst designs.

Xuefeng Li, Haoke Qiu, Hanwen Pei et al. · 0 citations
Jul 2026

AI-Guided High-Throughput and Uncertainty-Aware Discovery of Stable Na-Ion Cathodes

The rapid advancement of autonomous materials research requires data-efficient frameworks that can integrate artificial intelligence (AI) with experimental knowledge and human intuition. In the context of sodium-ion batteries (SIBs), where compositional and processing complexity spans an immense design space, conventional high-throughput experiments still face bottlenecks in data utilization and decision-making efficiency. To address this challenge, we developed an AI-guided, uncertainty-aware workflow that couples high-throughput synthesis and characterization [1-3] with machine learning (ML) surrogate modeling and probabilistic deep learning. Specifically, we established Monte Carlo (MC) Dropout–based multilayer perceptron (MLP) models to predict key electrochemical metrics such as initial capacity and long-term retention from high-dimensional descriptors encompassing composition, structure, and processing features. The MC-MLP provides not only accurate property predictions but also calibrated uncertainty estimates that quantify model confidence and guide active learning. These models are integrated with gradient boosting regression (GBR) surrogate mappings that link intermediate feature spaces to improve physical interpretability and reduce data sparsity. Tail-focused loss functions and temperature-balanced sampling further enhance sensitivity toward high-performance outliers, accelerating the identification of promising chemistries within limited datasets. By combining these ML strategies with automated synthesis and characterization workflows, we demonstrate an iterative human-AI loop: experimental data continuously refines the surrogate landscape, while uncertainty-driven candidate selection prioritizes new experiments. This approach achieved >3× improvement in data efficiency and enabled Pareto-optimized discovery of Na–Ni–Mn–based layered cathodes with enhanced air stability and electrochemical performance. Moreover, the framework is transferable across chemical systems, synthesis conditions, and multidimensional battery characterization data, enabling mechanistic understanding and accelerated discovery of novel electrode materials. Overall, this work highlights how integrating MC-Dropout deep learning, surrogate modeling, and domain knowledge from real lab-generated data can transform high-throughput experimentation into a self-improving, autonomous discovery framework—paving the way toward truly intelligent, data-efficient materials research. References [1] Jia, Shipeng, et al. "High-throughput design of Na–Fe–Mn–O cathodes for Na-ion batteries." Journal of Materials Chemistry A 10.1 (2022): 251-265. [2] Jia, Shipeng, et al. "Chemical speed dating: the impact of 52 dopants in Na–Mn–O cathodes." Chemistry of Materials 34.24 (2022): 11047-11061. [3] Jia, Shipeng, et al. "Stabilization of Na‐Ion Cathode Surfaces: Combinatorial Experiments with Insights from Machine Learning Models." Advanced Energy and Sustainability Research (2024): 2400051

Shipeng Jia, I. Abate · 0 citations
Open access Jul 2026

Distillation enables scalable high-fidelity virtual screening across ultra-large chemical libraries

Accurate virtual screening of ultra-large chemical libraries remains challenging. Existing approaches rely on lower-fidelity scoring functions or sampling-based strategies that can limit predictive accuracy and bias the exploration of chemical space. Here, we present FastBindRank, a distillation-based framework that transfers the predictive power of the structure-based model Boltz-2 into an efficient sequence-based surrogate. Trained on ∼1% of the 122-million-compound PubChem library, FastBindRank enables high-fidelity screening at scale. Applied to histone deacetylase 11 (HDAC11), FastBindRank substantially enriched high-confidence binders relative to the background chemical space. The lightweight model captured structural patterns associated with predicted binding, revealing structural determinants of binding. Under a comparable computational budget, FastBindRank achieved a 74-fold increase in hit rate and over a 30-fold increase in discovery yield over direct subset-based screening. Experimental validation confirmed the activity of two novel compounds. These results establish distillation as a practical strategy for scalable, high-fidelity virtual screening of ultra-large chemical libraries.

J. Dai, Yueyue Wang, N. Shan et al. · 0 citations
Preprint Jul 2026

Stoichiometric cluster learning for few-shot property prediction of multi-ionic integrated energetic materials

It is shown how pretrained machine-learned interatomic potentials (MLIPs) can bypass full crystal-structure prediction and support pre-synthesis screening from stoichiometric ionic clusters using multi-ionic integrated explosives (MIXs) as a synthesis-facing example.

Ming-Yu Guo, W. Zou, Yu Shang et al. · 0 citations
Jul 2026

(Invited) Constructing Multi-Scale and Flexible Effective Features to Accelerate Data-Driven Materials Design

Accelerating data-driven materials design is critical for advancing sustainable energy, optoelectronics, and catalysis, yet traditional approaches suffer from computational inefficiency, poor generalization, or inadequate capture of complex structure-property relationships, highlighting the urgent need for systematic and effective feature engineering. Herein, we present a series of our recent work to address this challenge. We proposed the LESS classifier, utilizing bond orientational order (BOO) parameters as lightweight features for crystal structure classification, enabling efficient recognition of mono/binary/amorphous and other crystal structures with over 98.8% accuracy and low retraining cost for new phases. Building on the LESS framework, we further developed luMOD, a 24-dimensional universal descriptor integrating convoluted BOO parameters and neighbor type encoding, delivering superior performance in multispecies systems (e.g., perovskites, olivines) with minimal computational overhead. Harnessing the robust encoding capacity of large language models (LLMs), we introduced EvoMD-LLM, which abstracts MD trajectories into macro-symbolic sequences to model species-level reaction dynamics, bridging static linguistic knowledge with dynamic temporal evolution of chemical systems. Additionally, we proposed a transferable bandgap prediction framework for perovskites, integrating ChemGPT-derived atomic embeddings with global attention graph networks to capture organic-inorganic synergies and enable reliable extrapolation to unseen compositions. Collectively, these studies advance feature engineering across static/dynamic, single/multispecies, and equilibrium/reactive systems, unifying interpretability, efficiency, and transferability. They aim to nhance the reliability and scalability of data-driven materials discovery, propelling autonomous functional materials design.

Yanming Wang · 0 citations