Multimodal large language models (MLLMs) have made strong progress on visual question answering and image captioning, yet they still produce fluent claims about objects, attributes, or relations that are not grounded in the image. Many remedies either modify decoding at test time, which adds latency, or fine tune with preferences such as DPO variants, which teach which answer is preferred but not when the model's own answer is unreliable. We argue that calibrated self assessment is the missing signal. We introduce Savor, a training framework that (i) augments the output schema with token and answer confidence, (ii) optimises the policy with a Group Relative Policy Optimisation (GRPO) objective that penalises calibration error and poor abstention decisions, and (iii) uses the learned confidence at inference time to revisit visual evidence only when the model is uncertain. Experiments on POPE, HallusionBench, AMBER and MMHal-Bench across two recent backbones (InternVL3-8B and Qwen3-VL-8B) show that Savor reduces hallucination while preserving general capability on MME and MMBench, with lower Expected Calibration Error than DPO and decoding baselines.
Zian Ding, Zi-Lin Zhao, Ying-Jie He et al.· 0 citations
The monitoring of mental health states using electroencephalogram (EEG) signals has gained increasing attention due to its non-invasive nature for psychological disorders. Large Language Models (LLMs) and Explainable Artificial Intelligence (XAI) have been utilized in advancing the intelligence and interpretability of EEG analysis. However, existing methods face critical bottlenecks, including the fundamental modal gap, high computational costs, and poor global consistency. The limitation of rigid classification tasks without supporting clinical reasoning and natural language interaction. In this study, we propose a collaborative explainable AI framework for EEG mental health monitoring with constrained question-and-answer (QA) tuned LLM alignment, which builds a smooth transformation path from raw EEG signals to evidence, and constructs a structured QA dataset for the instruction fine-tuning of LLMs. The central objective of this work is not simply to maximize EEG classification accuracy, but to develop an evidence-grounded alignment and explanation framework that connects EEG-derived physiological evidence with QA-based LLM reasoning. Furthermore, this work designs a transparent collaborative XAI mechanism that embeds interpretable EEG feature information as prior knowledge directly into the QA generation process of the LLM, and develops a multi-level interpretable pipeline combining attention heatmap analysis and decision tree surrogate modeling to achieve precise alignment between LLM internal reasoning and EEG neurophysiological patterns. The proposed framework addresses the limitations of traditional rigid EEG classification tasks, promotes the XAI paradigm shift from high-cost post-hoc explanations to transparent embedded explanations, and enables robust clinical reasoning and natural language interaction based on EEG signals. Experimental results on a benchmark EEG mental state dataset demonstrate that the proposed framework stably captures neurophysiological characteristics corresponding to different mental states, and effectively improves decision transparency and clinical credibility of EEG-based mental health monitoring systems. In this setting, classification performance is treated as one evaluation aspect, while the primary contribution lies in constrained evidence-grounded alignment and QA-based LLM explainability. This advancement provides an initial feasibility study of real-time, scalable, and trustworthy intelligent EEG-based mental health analysis.
Zian Ding, Fusen Guo, Bonan Zhang et al.· Journal of King Saud Univers...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.