Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet reliability requires more than accuracy: a model must also use the ordinal decision scale supplied by the user faithfully. We analyze JEV~1.13 and three open KEV models...
Tian-Xiang Gao, Jin-Zhe Li, Zhiyuan Li et al.· 0 citations
Multi-turn interaction creates a feedback process in which an LLM's previous responses become context for later behavior. Prior work shows substantial multi-turn degradation and that assistant-generated history can affect later behavior. However, it remains unclear how these effects manifest across models, tasks, turns...
Jinnan Li, Zheren Fu, Yue Wang et al.· 0 citations
This work proposes MMPCBench, a comprehensive framework for evaluating MLLMs' proactive critique competence, which adopts a hierarchical evaluation protocol to measure models' error detection, diagnosis and resolution performance, and applies alignment-aware metrics to assess the coherence between internal reasoning an...
Jin-Zhe Li, Geng-Xu Li, Jinnan Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.