Resource-efficient machine learning pathways for molecular discovery in organic photovoltaics
Abstract
This thesis presents a data−driven approach to accelerate the discovery and design of materials for organic photovoltaic devices, while reducing the usage of computational resources. The core of this work was the development of machine learning models for property prediction of organic photovoltaics and to implement a generative artificial intelligent model for molecular discovery. Fundamental to this idea was the construction of a curated database containing pairs of donor and acceptor molecules relevant to organic solar cells, referred to as the Pairs of Donors and Acceptors database. Additionally, the database includes information on frontier molecular orbitals photovoltaic performance parameters such as power conversion efficiency, charge carrier mobilities, and molecular representations compatible with widely adopted cheminformatics formats such as the Simplified Molecular Input Line Entry System. This database is intended to serve as the basis for building efficient molecular representations and for training machine learning models to predict chemical properties. To improve memory efficiency and reduce training time, a novel and fully reversible embedding technique for categorical data, applicable to one−hot encodings and Morgan Fingerprints, was introduced. These embedded representations, named embedded One Hot Encoding and embedded Morgan Fingerprints respectively, were evaluated in both generative and predictive contexts, demonstrating performance comparable to, or even better than, traditional cheminformatics encodings for machine learning applications. Subsequently, deep learning models were trained using embedded Morgan Fingerprints to predict frontier molecular orbital energy levels with high accuracy, including a transfer learning strategy validated with external databases beyond the in−house constructed Pairs of Donors and Acceptors database. The Pairs of Donors and Acceptors database was used to train generative models, from which, after a filtering process, over 9.4 × 104 donor molecules and more than 2.1 × 105acceptor molecules were identified. The prediction of their frontier orbital energy levels was carried out using the machine learning models pretrained with frontier molecular orbitals energy values, enabling the proposal of novel donor−acceptor molecular pairs potentially relevant to organic photovoltaic applications. Finally, an outlook is presented focusing on the development of power conversion efficiency prediction models, along with computational strategies to validate promising candidates using density functional theory calculations and structure−energy alignment. The ultimate goal is to enable accelerated exploration of organic materials for solar energy, contributing to the advancement of more efficient renewable energy technologies.