Multimodal Large Language Models (MLLMs) suffer from hallucinations, creating a critical need for Uncertainty Quantification (UQ) to ensure reliable deployment. However, existing approaches struggle to detect uncertainty caused by superficial associations, especially when the query-relevant signal is weak. We mainly at...
Hao-Yang Luo, Lin-Wei Tao, Jie Gui et al.· 0 citations
Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduces these costs by rendering history as images, but the resulting modality shift creates a marked capability gap. Through controlled evaluations of history recovery, matc...
Cheng Fan, Junyi Zhou, Tingzhang Luo et al.· 1 citation
RoRA is a training-free framework that casts visual token pruning as role-oriented regional evidence allocation, and consistently outperforms strong training-free baselines across LLaVA and Qwen-VL families, retaining most of the unpruned accuracy even at aggressive pruning ratios.
Qiyanhui Lu, Han Wu, Rong-Jia Xu et al.· 1 citation
Trajectory-Aware Commit Gating (TACG), a training-free gate-level decoder that anchors token identities to the base posterior and uses trajectory-aware signals only to decide whether the current proposal is ready to commit, is proposed.
Chengcheng Wang, Tingzhang Luo, Wenhao Li et al.· arXiv.org· 3 citations
CROSS is proposed, a tightly integrated paradigm for RRSIS that achieves state-of-the-art performance and maintains precise localization even under severe spatial description perturbations, standing as a robust new paradigm for RRSIS.
Tingzhang Luo, Ruizhong Liu, Yi-Chao Liu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.