From QSAR to deep learning: an interpretable comprehensive pipeline with a read-across approach for mutagenicity prediction via the Enalos Cloud Platform
In silico NAMs, including computational approaches, can contribute to the development of novel Safe and Sustainable by Design chemicals and substances by identifying potentially hazardous ones at an early stage by identifying potentially hazardous ones at an early stage.
Abstract
Assessing the mutagenicity of chemical compounds is essential for ensuring their safe handling and use, thereby minimizing potential health risks. New approach methodologies (NAMs), including computational approaches, provide non-animal alternatives for testing novel materials and chemicals. This study highlights the potential of in silico NAMs, which can contribute to the development of novel Safe and Sustainable by Design (SSbD) chemicals and substances by identifying potentially hazardous ones at an early stage. Emphasis is given to the mutagenicity prediction based on data from the Ames (bacterial gene mutation) test curating them to consider stereo-specific input whenever necessary. A consensus strategy integrating different chemical representations (molecular fingerprints, 2D and 3D descriptors and molecular graphs) and modelling methods, i.e., Quantitative Structure–Activity Relationship (QSAR), read-across and deep learning models, is employed to predict the mutagenic profile of chemical compounds. In this course, a XGBoost model is developed based on molecular descriptors and a graph convolutional neural networks model to classify compounds as mutagens and non-mutagens based on the Ames test data. The devised read-across methodology is based on a guided-k-Nearest Neighbours scheme (guided-kNN) where two different molecular representations (molecular fingerprints and descriptors) are considered for neighbour selection and predictions generation. The mutagenicity predictions from the three models are integrated in a majority voting scheme to enhance the overall predictive accuracy (83% in external validation) and reduce individual model biases. Interpretation of the descriptors involved in prediction is performed through explainable AI (XAI) methods to provide insight to the mutagenicity mechanism. To enhance the interpretability of the XAI-derived insights and reinforce user confidence in the models' predictions, the involved descriptors are mapped to key events leading to mutations within the Adverse Outcome Pathway (AOP) networks. Apart from the development of reliable and interpretable mutagenicity models, emphasis is given on delivering a pipeline for the generation of 3D descriptors that can be used as the basis for future cheminformatics models. To support transparency and reproducibility of the results of our work, the curated mutagenicity dataset used for modelling is disseminated through the ChemPharos database (https://db.chempharos.eu/datasets/Datasets.zul?datasetID=ds18), the modelling steps are documented following the standardized Modelling Data (MODA) guidelines and the consensus model is freely available via the Enalos Cloud platform (https://www.enaloscloud.novamechanics.com/insight/polis/), to facilitate virtual screening of novel compounds.
The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.
Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson· EUROMICRO Conference on Soft...· 64 citations· ⚡6
The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.
Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al.· arXiv.org· 62 citations· ⚡3
The use of large language models to automatically improve the user story quality in Austrian Post Group IT agile teams is explored, with a reference model for an Autonomous LLM-based Agent System developed and implemented at the company.
Zheying Zhang, M. Rayhan, Tomas Herda et al.· International Conference on...· 48 citations· ⚡4
This paper introduces a novel multi-AI-agent system designed to fully automate SLRs, and demonstrates how it substantially reduces the time and effort traditionally required for SLRs while maintaining comprehensiveness and precision.
Abdul Malik Sami, Z. Rasheed, Kai-Kristian Kemell et al.· arXiv.org· 44 citations· ⚡2
The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners, and future improvements focus on enhancing multilingual performance and integrating continuous expert feedback.
Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al.· arXiv.org· 41 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.