A clinician-in-the-loop benchmark for evaluating whether large language models can generate evidence-grounded Brief Hierarchical Taxonomy of Psychopathology item profiles from passive sensing, ecological momentary assessment (EMA), and questionnaire evidence is introduced.
Abstract
Mobile and wearable sensing enables longitudinal observation of behavior, yet translating these signals into meaningful mental health constructs remains difficult. We introduce a clinician-in-the-loop benchmark for evaluating whether large language models (LLMs) can generate evidence-grounded Brief Hierarchical Taxonomy of Psychopathology (B-HiTOP) item profiles from passive sensing, ecological momentary assessment (EMA), and questionnaire evidence. Using the Generalization of Longitudinal Behavior Modeling (GLOBEM) dataset, we construct 14,592 participant-day instances and align multimodal evidence to 29 B-HiTOP items across five spectra. Since GLOBEM lacks B-HiTOP responses, we evaluate evidence compatibility (C) rather than diagnostic accuracy, separating substantive predictions from abstentions when evidence is insufficient for item-level scoring. Two-stage prediction improves C for EMA and questionnaire evidence, but reduces C under passive sensing and combined evidence and produces more conservative score distributions across models, spectra, and evidence settings. Overall, semantic abstraction helps organize heterogeneous self-report evidence while becoming an information bottleneck for indirect behavioral sensing signals.
Brief multimodal wearable assessment as an objective complement to caregiver-reported screening supports brief digital phenotyping as an objective complement to caregiver-reported screening.
B. Loftness, J. Cohen, D. Kairamkonda et al.· medRxiv· 0 citations
WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.
Ji Soo Lee, Xi-Lun Chen, Pierce Chuang et al.· 0 citations
Mental health conditions are widespread, yet scalable assessment remains limited, as clinical interviews are time-intensive and self-report screeners are episodic and prone to bias, motivating passive, data-driven approaches. This study investigates whether short-term physical activity patterns captured by consumer wea...
Guan-Nan Liu, Jae Yoo, Gao-Jian Huang et al.· Proceedings of the Internati...· 0 citations
BALMS is the first systematic benchmark of LLM-based agentic systems for longitudinal mental health sensing and highlights the need for longitudinal mental health agents that selectively retrieve history, ground temporal evidence, and reason over interpretable behavioral features.
Y. Wu, Arvind Pillai, Yu-Liang Chen et al.· 0 citations
Repeated multimodal experiments in human–computer interaction (HCI) generate hierarchically and temporally structured data, making analyses that treat sensor-derived observations as independent vulnerable to invalid inference. The primary contribution of this study is a reproducible, dependence-aware multilevel framewo...
Sara Morgado-Nunes, C. Patino-Alonso· Multimodal Technologies and...· 0 citations