Vision-language models (VLMs) have demonstrated excellent performance in visual tasks, but their visual causal reasoning capabilities still lack reliable evaluation. Existing evaluations struggle to distinguish whether a model is performing causal reasoning based on visual evidence or relying on statistical correlation...
Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet reliability requires more than accuracy: a model must also use the ordinal decision scale supplied by the user faithfully. We analyze JEV~1.13 and three open KEV models...
Tian-Xiang Gao, Jin-Zhe Li, Zhiyuan Li et al.· 0 citations
Multi-turn interaction creates a feedback process in which an LLM's previous responses become context for later behavior. Prior work shows substantial multi-turn degradation and that assistant-generated history can affect later behavior. However, it remains unclear how these effects manifest across models, tasks, turns...
Jinnan Li, Zheren Fu, Yue Wang et al.· 0 citations
This work proposes MMPCBench, a comprehensive framework for evaluating MLLMs' proactive critique competence, which adopts a hierarchical evaluation protocol to measure models' error detection, diagnosis and resolution performance, and applies alignment-aware metrics to assess the coherence between internal reasoning an...
Jin-Zhe Li, Geng-Xu Li, Jinnan Li et al.· 0 citations
Experiments with representative closed-source and open-source MLLMs show that OCR-grounded meta-reasoning remains far from saturated: models struggle with visible-rule application and layout-sensitive inference, while process-compliant rationales can accompany incorrect final answers under exact-match evaluation.
Geng-Xu Li, Yuan Wu, Yi Chang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.