Skip to content

Transparency and Explainability in Human Factors — A Systematic Review of Usability Assessment Practices for AI Medical Devices

Sep 2026 · Vol 15, pp. 62 - 66 · 0 citations · 5 references

TL;DR

Results identify a "symmetry of modality": qualitative interviews correlate with written text explanations, while Think-Aloud protocols better assess cognitively demanding tools like SHAP values.

View source

Similar papers

Book Open access Jul 2026

A Multidimensional Approach to Usability Engineering, Clinical Performance and Human-AI Interaction in AI-Enabled Medical Devices

Artificial Intelligence-Enabled Medical Devices (AIeMD) promise to revolutionize healthcare, yet their safe adoption relies on effective Human-AI Interaction (HAAI) design and validation. Established usability engineering standards and guidances, including IEC 62366-1 and FDA frameworks, fail to address the novel sociotechnical risks of "black-box" systems, including automation bias and the misalignment of clinician mental models. This doctoral project directly addresses this gap, with the primary goal of developing a tailored human factor/usability engineering framework specifically for AIeMD. The project aims to establish robust methodologies for evaluating transparency and trust, integrating human factors and clinical performance into a multidimensional validation pipeline. Ultimately, this work will provide the evaluative tools necessary to move from subjective satisfaction to safety-critical, risk-based validation, ensuring that AI-enabled health solutions are clinically reliable, transparent, and compliant

Mariana de Oliveira · 0 citations
Book Open access Jul 2026

Human-AI Interaction in Healthcare - A Manifesto for Improved Usability Evaluation

Traditional usability assessments and questionnaires, such as the System Usability Scale (SUS), were designed for deterministic systems with predictable, linear outputs. However, AI-enabled medical devices are inherently probabilistic and co-evolve with the user through repeated interaction, rendering traditional usability assessments insufficient for guaranteeing the long-term safety in the use of high-risk probabilistic systems. Current literature reveals a striking absence of longitudinal studies, creating significant methodological blind spots regarding how trust calibrates over time and whether automation bias intensifies with habitual use. In this position paper, we present a manifesto for a longitudinal, three-fold methodological pivot in health human-AI interaction. We propose moving beyond static satisfaction metrics towards relational metrics —Longitudinal Trust Calibration (LTC), Automation Bias Drift (ABD), and Error Recovery Velocity (ERV)—that track the maturity and resilience of the human-AI partnership. This framework provides an actionable path toward a safety-in-use paradigm that acknowledges the temporal, dynamic nature of high-risk health AI.

Mariana de Oliveira, Célia F. Cruz, Nuno Matela · 0 citations
Conference Open access 2026

Explainable AI in Telemedicine: A Scoping Review of Clinical Usability Gaps and Research Directions for GP-Centred Design

: Remote healthcare services have grown significantly since the COVID-19 pandemic, with AI increasingly deployed in telemedicine for diagnosis support, triage, and monitoring. While Explainable AI (XAI) methods have been proposed to address trust and transparency concerns, their adoption in clinical practice remains limited — particularly in general practice settings where time constraints, diagnostic breadth, and patient-facing communication impose unique demands on explanation design. This paper presents a scoping review of XAI applications in telemedicine, examining 20 studies published between 2022 and 2026 to characterise the methods used, their evaluation approaches, and gaps between technical explainability and clinical usability. SHAP was the dominant method across the corpus, yet only three papers operated in a telemedicine context, and none evaluated XAI with GPs or in primary care. Empirical evaluation with clinical users was rare, small-scale, and methodologically inconsistent, with a median sample size of 21 participants across five empirical studies. We identify systematic gaps in telemedicine context, GP user focus, and HCI integration, and conclude with research directions for human-centred XAI in telemedicine, derived from identified gaps.

Ana Vesic, M. Zivkovic · 0 citations
Open access Jul 2026

Cognitive Friction in Clinical Decision Support: A Comparative Study of Judicial and Adjunct Human–AI Interaction Protocols

Artificial intelligence is increasingly used to support clinical decision making, yet concerns remain regarding algorithmic aversion, automation bias and the preservation of meaningful human oversight; while explainable AI aims to improve transparency, less attention has been devoted to the design of human–AI interaction protocols. This study investigates Frictional AI, an interaction paradigm that introduces cognitive friction to encourage critical engagement with AI recommendations. First, semi-structured interviews were conducted with a legal expert and a psychologist and analyzed through thematic analysis to identify legal, ethical, and cognitive requirements for AI-assisted decision support. Second, a user study involving 96 medical residents compared three interaction protocols: a conventional explainable AI-first design (XAI) and two friction-based protocols, namely a judicial protocol based on juxtaposed explanations (Judicial AI, JAI) and an adjunct protocol requiring an initial unsupported decision before AI exposure (AAI). Diagnostic accuracy and confidence, perceived usefulness, completion time, and reliance patterns were evaluated. The interviews highlighted the importance of human-centered explanations, contrastive reasoning, preservation of professional responsibility, and the role of user studies in evaluating human–AI interaction. The quantitative results showed that none of the AI-assisted conditions improved diagnostic accuracy relative to the no-support baseline. However, JAI achieved performance comparable to the baseline, outperforming XAI and AAI, and exhibited the lowest level of over-reliance. Overall findings suggest that the effectiveness of decision-support systems depends not only on model performance and explanation quality but also on interaction design. In conclusion, while preserving diagnostic performance, judicial protocols showed promise in mitigating automation bias and promoting active cognitive engagement in clinical decision support.

Samuele Pe, Laura Bergomi, G. Nicora et al. · 0 citations
Review Open access Aug 2026

Validation of an AI-powered mobile application for personalizing medical note explanations: a mixed-methods evaluation

Nearly half of adults struggle to understand written health information, making medical communication a persistent barrier to effective care. While artificial intelligence has potential to improve health communication, few patient-facing tools have undergone systematic validation for personalized medical explanation. Patiently AI is a mobile application designed to clarify clinician-authored medical notes using large language models with audience-specific adaptations (child, teenager, adult, carer) and tone variations (friendly, informative, reassuring). A three-phase mixed-methods evaluation was conducted: (1) computational readability analysis of 210 AI-generated explanations using established metrics; (2) expert review by 15 healthcare professionals assessing medical accuracy, safety, and communication quality; and (3) a patient survey of 54 participants evaluating preferences, comprehension, and acceptance. AI-generated explanations demonstrated consistent improvements in readability, with mean Flesch–Kincaid Grade Level decreasing by 2.96 levels (10.57–7.61), Flesch Reading Ease increasing by 31.9 points (37.7–69.6), and Gunning Fog Index decreasing by 4.09 points (14.5–10.4); all improvements were statistically significant (all P  ≤ 0.002). Readability gains were greatest for younger audiences (child: 4.25 grade-level reduction; adult: 1.80). Expert reviewers rated outputs highly for medical accuracy (4.49 ± 0.83/5), clarity (4.53 ± 0.77/5), and trustworthiness (4.37 ± 0.90/5), with 87.3% assessed as clinically safe. Inter-rater agreement across the 15 reviewers was substantial (Gwet's AC1 = 0.72 for safety assessments). Among patients, 70.0% of responses preferred AI-generated explanations ( P  < 0.001), with 98.1% comprehension accuracy and high ratings for clarity (4.58 ± 0.65/5) and confidence in care (4.19 ± 0.85/5). Overall, 70.4% indicated a likelihood of using the application. This mixed-methods evaluation suggests that a deliberately constrained, language-focused AI system can improve the accessibility of medical notes while preserving clinical accuracy and safety. Patiently AI demonstrates a scalable approach to supporting health literacy and patient engagement without extending into clinical interpretation.

Nicholas Lamb · 0 citations
Open access Aug 2026

Lessons from the field: usability (engineering) in regulated healthcare

Abstract Regulated healthcare settings continue to face gaps between established HCI methodology, user-centered design principles, and everyday clinical practice. Although standards and regulatory frameworks mandate usability engineering, their integration into regulated clinical settings often remains inconsistent, compliance-driven, and detached from real-world workflows. To better understand and address this disconnect, we report insights from applied work in clinical studies, hospital collaborations, and industry partnerships. Our aim is to identify methodological and organizational challenges in assessing digital health technologies and discuss potential strategies to mitigate them. While usability evaluations are typically conducted in accordance with international standards, our observations reveal systemic problems: hospital stakeholders are rarely involved in early design stages, procurement decisions are frequently made without structured user testing, and patients are often excluded from technology selection processes. We argue that usability engineering should be a formative, safety-critical discipline integrated throughout the product lifecycle, including aftermarket deployment, and not merely a compliance task. Drawing on case study reflections and practitioner experience, we articulate actionable lessons and practical strategies for integrating usability engineering more effectively into regulated healthcare environments. By bridging methodological, regulatory, and organizational divides, the paper contributes to guidance for aligning human-centered design with the needs of complex healthcare systems.

Preetha Moorthy, T. Nagel, Eva Hornecker · 0 citations

Related blog posts