Human-centred approaches for generating and evaluating explainable AI across modalities
Abstract
Since work on this thesis began in October 2022, artificial intelligence has become increasingly prevalent across everyday life, scientific research, and the global economy. ChatGPT was released the following month and, by 2026, had more than one billion weekly users. In 2024, computer scientists were among the recipients of both the Nobel Prize in Physics and the Nobel Prize in Chemistry for work closely connected to modern artificial intelligence, and in 2025 NVIDIA became the first company to reach a market capitalisation of $5 trillion. This rapid expansion increases the importance of explainable artificial intelligence (XAI), which develops methods for making AI behaviour more understandable to people. My research advances human-centred XAI by developing and evaluating explanation methods that help people understand model behaviour, judge when to rely on AI systems, and improve those systems. This thesis addresses three challenges in explanation: identifying recurring behaviour across image explanations, communicating complex explanations through natural-language narratives, and understanding the risks created by explanations. The first part develops methods for aggregating local counterfactual and saliency explanations to reveal recurring behaviour in image classifiers without requiring manual inspection of many individual images. The second investigates how complex explanations can be communicated through natural-language narratives. The third examines how explanation systems can be used to mislead users about model behaviour and how this risk can be reduced.