(Invited) Constructing Multi-Scale and Flexible Effective Features to Accelerate Data-Driven Materials Design
Abstract
Accelerating data-driven materials design is critical for advancing sustainable energy, optoelectronics, and catalysis, yet traditional approaches suffer from computational inefficiency, poor generalization, or inadequate capture of complex structure-property relationships, highlighting the urgent need for systematic and effective feature engineering. Herein, we present a series of our recent work to address this challenge. We proposed the LESS classifier, utilizing bond orientational order (BOO) parameters as lightweight features for crystal structure classification, enabling efficient recognition of mono/binary/amorphous and other crystal structures with over 98.8% accuracy and low retraining cost for new phases. Building on the LESS framework, we further developed luMOD, a 24-dimensional universal descriptor integrating convoluted BOO parameters and neighbor type encoding, delivering superior performance in multispecies systems (e.g., perovskites, olivines) with minimal computational overhead. Harnessing the robust encoding capacity of large language models (LLMs), we introduced EvoMD-LLM, which abstracts MD trajectories into macro-symbolic sequences to model species-level reaction dynamics, bridging static linguistic knowledge with dynamic temporal evolution of chemical systems. Additionally, we proposed a transferable bandgap prediction framework for perovskites, integrating ChemGPT-derived atomic embeddings with global attention graph networks to capture organic-inorganic synergies and enable reliable extrapolation to unseen compositions. Collectively, these studies advance feature engineering across static/dynamic, single/multispecies, and equilibrium/reactive systems, unifying interpretability, efficiency, and transferability. They aim to nhance the reliability and scalability of data-driven materials discovery, propelling autonomous functional materials design.