It is shown that hallucinations often accumulate over long contexts, through self-reinforcing dialogue history, and models are particularly vulnerable to questions requiring premise rejection or refusal, which highlights CEDI as a step toward realistic, systematic, and ecologically valid assessments of MLLMs'capabilities.
Abstract
Multi-modal Large Language Models (MLLMs) have made substantial advances on benchmarks, yet their real-world effectiveness remains uncertain. This gap stems from the fundamental misalignment between benchmarks in controlled, static settings and the dynamic, interactive, and contextualized nature of real-world applications. To bridge this gap, we propose CEDI (Contextualized Evaluations of MLLMs through Dynamic, multi-round Interactions), a framework that recasts evaluation as a three-party interaction between an evaluatee model, an automated examiner, and a grader. The examiner conducts multi-turn, semi-structured conversation guided by a graph-based representation of the task. By navigating state-space transitions, CEDI deploys diverse strategies, from clarification requests to adversarial probes, to elicit performance evidence. We apply CEDI to visual hallucinations. Empirical results across multiple models, diverse settings, datasets, and domains show that contextualized, interactive evaluations reveal not only significantly more hallucinations than conventional static evaluation but also ones that more closely resemble those arising in practical use cases. We further show that hallucinations often accumulate over long contexts, through self-reinforcing dialogue history, and models are particularly vulnerable to questions requiring premise rejection or refusal. Together, these findings highlight CEDI as a step toward realistic, systematic, and ecologically valid assessments of MLLMs'capabilities. Code is available at github.com/williamium3000/cedi.
The proposed Multi-turn Multimodal Interactive Dialogue (MMID) Benchmark enables comprehensive evaluation of the Perception, Memorization, and Reasoning abilities of MLLMs, and reveals MLLMs perform well with text-based input but degrade with images, requiring improved leverage fine-grained visual cues.
Seulgi Kim, Juoh Sun, Sumin Kim et al.· Proceedings of the 32nd ACM...· 0 citations
This work introduces DocHop, a benchmark for integrated chart--context reasoning in document-style images and constructs DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, to enable systematic evaluation.
Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al.· 1 citation
This work presents MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs, and introduces an LLM-as-a-Judge metric to assess the correctness of model reasoning.
A novel knowledge graph-based evaluation framework is proposed introducing Semantic Structural Similarity for KGs (S3KG), a hybrid similarity measure integrating structural and semantic similarity into a continuous evaluation score, alongside a diagnostic framework for categorizing reasoning errors.
Subavarshana Arumugam, Mamta Nallaretnam, K. Wickramasinghe et al.· 0 citations
An integrated conceptual frame-work that couples attention- and perturbation-based explainability with lightweight hallucination-detection signals and token-efficient inference strategies is presented, and a set of cross-cutting consistency metrics are instrumented with a set of cross-cutting consistency metrics.
Sakshi Parate, Shreyans Sanyal· Advanced International Journ...· 0 citations
AutoDavis is introduced, a first-of-its-kind automatic and dynamic evaluation protocol that enables on-demand benchmarking of LVLMs across specific capability dimensions and shows effectiveness and reliability, offering a new paradigm for dynamic benchmarking of multimodal intelligence.
Han Bao, Yue Huang, Yan-Bo Wang et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.