Skip to content

A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series

Jul 2026 · arXiv.org · Vol abs/2607.25947 · 1 citation · 28 references
Computer Science

TL;DR

ClinPRISM, a cost-effective multimodal LLM reasoning framework for question answering over ICTS data, is proposed and achieves state-of-the-art performance on the held-out evaluation benchmark while using only 16 time-series tokens and achieving an average inference latency of 0.15 seconds per question.

Abstract

Question answering (QA) over irregular clinical time series (ICTS) plays a pivotal role in a wide range of healthcare applications. Although recent multimodal time-series large language models (LLMs) have shown considerable promise in general-purpose time-series QA, they remain poorly equipped to model the sparsity, asynchrony, and irregular sampling patterns of clinical observations. To fill this gap, we propose ClinPRISM, a cost-effective multimodal LLM reasoning framework for question answering over ICTS data. First, we devise an irregularity-aware multi-scale encoder to capture sparse clinical evidence at diverse temporal scales. Then, we propose a temporal evidence distiller to integrate representations across these scales and compress them into a small number of LLM-compatible tokens. Moreover, we introduce a progressive alignment strategy that sequentially aligns the irregular trajectories with the LLM's textual embedding space. To facilitate training, we construct 30,000 clinical time series paired with multi-scale descriptions, together with 41,000 instruction-tuning instances spanning 11 tasks. Using a 4-billion-parameter LLM backbone, ClinPRISM achieves state-of-the-art performance on the held-out evaluation benchmark while using only 16 time-series tokens and achieving an average inference latency of 0.15 seconds per question.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain

MMTClinic is presented, a benchmark designed to evaluate large language models (LLMs) on complex reasoning and question-answering tasks involving clinical time-series and reveals notable differences in model performance across tasks, languages, and modalities, highlighting current limitations in clinical reasoning capa...

Sourav Malakar, Harshit Nigam, Akash Ghosh et al. · 0 citations
Jul 2026

LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA

LoMeVQA, a comprehensive benchmark consisting of 206K longitudinal visual question answering (VQA) pairs for temporal medical image analysis, is proposed and MedLong-8B, which achieves state-of-the-art performance across all tasks is introduced.

Zhilin Wu, Zhangkai Ni, Cheng Yang et al. · 0 citations
Open access Aug 2026

CARE-LLM-GRAPH: Confidence Aware LLM integrated Multimodal Architecture for clinical Recommendation

A new confidence-aware hybrid design, CARE-LLM-GRAPH, which combines large language models (LLMs) to perform clinical reasoning, multimodal deep learning to analyze medical images, and population-aware graph intelligence to provide cohort-level information is presented.

Unknown authors · 0 citations
Preprint Aug 2026

MedReaMM: Evaluating Large Multimodal Models on Expert-Level Clinical Diagnostic Synthesis

This work introduces MedReaMM, a benchmark specifically designed to evaluate models'ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm.

Lai Wei, Yu-Chao Chen, Zhenbiao Cao et al. · 0 citations
Open access Aug 2026

MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation

This work presents MTDiag, a large multi-turn diagnostic dialogue dataset constructed from three heterogeneous sources: DDXPlus, MIMIC-IV, and published case reports, and introduces and motivate clinical knowledge-grounded metrics for evaluating LLMs as diagnostic agents, beyond diagnostic accuracy, for the task of mul...

Pia Chouayfati, Alexander M. Fichtl, Miriam Anschütz et al. · 1 citation
#small language model Preprint Aug 2026

Future Querying: Can LLMs Serve as Implicit Medical World Models?

This work introduces future querying, a paradigm that probes whether large language models can function as implicit medical world models by evaluating their ability to answer time-indexed clinical queries about a patient's future, and shows that small, locally fine-tuned open-weight models can match or approach larger...

Siri Willems, James Butterworth, L. Goetschalckx et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.