Traditional usability assessments and questionnaires, such as the System Usability Scale (SUS), were designed for deterministic systems with predictable, linear outputs. However, AI-enabled medical devices are inherently probabilistic and co-evolve with the user through repeated interaction, rendering traditional usability assessments insufficient for guaranteeing the long-term safety in the use of high-risk probabilistic systems. Current literature reveals a striking absence of longitudinal studies, creating significant methodological blind spots regarding how trust calibrates over time and whether automation bias intensifies with habitual use. In this position paper, we present a manifesto for a longitudinal, three-fold methodological pivot in health human-AI interaction. We propose moving beyond static satisfaction metrics towards relational metrics —Longitudinal Trust Calibration (LTC), Automation Bias Drift (ABD), and Error Recovery Velocity (ERV)—that track the maturity and resilience of the human-AI partnership. This framework provides an actionable path toward a safety-in-use paradigm that acknowledges the temporal, dynamic nature of high-risk health AI.
Mariana de Oliveira, Célia F. Cruz, Nuno Matela· Information Hiding· 0 citations
Results identify a "symmetry of modality": qualitative interviews correlate with written text explanations, while Think-Aloud protocols better assess cognitively demanding tools like SHAP values.
M. A. D. De Oliveira, Constança Roquette, Nuno Matela et al.· 0 citations