A multimodal large language model (MLLM) council is developed that, given an image and its CBM explanation, produces an explanation quality score, and CBX-Bench, a public benchmark and leaderboard, provides a human-aligned, scalable evaluation of CBM explanations beyond accuracy and isolated qualitative examples.
Abstract
Concept Bottleneck Models (CBMs) are designed to make visual classification interpretable by expressing predictions through human-understandable concepts. Although interpretability is the central motivation for CBMs, they are still largely evaluated as predictive models by downstream classification accuracy, supplemented by isolated qualitative examples. This highlights a pressing need for quantitative measures, a challenge complicated by the infeasibility of ground-truth concept annotation at scale and the open nature of concept lists due to a lack of consensus. To fill this gap, we develop a multimodal large language model (MLLM) council that, given an image and its CBM explanation, produces an explanation quality score. To ground and validate the council, we first conduct a human study to establish a ground-truth reference for CBM explanation quality: for an image, annotators compare explanations from two of LF-CBM, VLG-CBM, and CBM-Suite and choose the more useful one, or mark them as equally good or equally bad, yielding 2700 judgments over 900 image-comparison items on CUB-200, ImageNet-100, and Places365. Against this human reference, our five-model council, consisting of open-weight MLLMs, recovers over 70% of strict human preference rankings, rising to 83% on items where human annotators unanimously agree. Building on this validated council, we introduce CBX-Bench, a public benchmark and leaderboard: authors of new CBMs can submit their model's explanations, and CBX-Bench scores them with the council and maintains dataset-level rankings of explanation quality. CBX-Bench thus provides a human-aligned, scalable evaluation of CBM explanations beyond accuracy and isolated qualitative examples. The benchmark is available at https://github.com/meric-karadag/cbx-bench.
Interpretable-by-design architectures, such as Concept-Bottleneck Models (CBMs), are essential for trustworthy AI under regulations like the EU AI Act. While classical CBMs rely on costly expert labels, recent automated methods use vision–language models for concept discovery. However, whether automation preserves inte...
Vincenzo Bevilacqua, A. Di Marino, A. Ciaramella et al.· Applied Sciences· 0 citations
This work introduces DocHop, a benchmark for integrated chart--context reasoning in document-style images and constructs DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, to enable systematic evaluation.
Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al.· 1 citation
This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and expli...
DiverValue-Bench is introduced, a population-aware benchmark for evaluating multi-dimensional value alignment across 74 countries/regions and it is shown that lightweight preference-based fine-tuning with Low-Rank Adaptation and Direct Preference Optimization substantially improves in-domain value alignment while yield...
Yao Liang, Dongcheng Zhao, Fei-Fei Zhao et al.· Proceedings of the Thirty-Fi...· 0 citations
A novel knowledge graph-based evaluation framework is proposed introducing Semantic Structural Similarity for KGs (S3KG), a hybrid similarity measure integrating structural and semantic similarity into a continuous evaluation score, alongside a diagnostic framework for categorizing reasoning errors.
Subavarshana Arumugam, Mamta Nallaretnam, K. Wickramasinghe et al.· 0 citations
This work presents an end-to-end explainable evaluation metric that fine-tunes a language model to identify omitted, extra, incorrect, and correct data units in a data-text pair, providing both fine-grained diagnostic feedback and an interpretable measure of alignment quality.
Kun Efimov-Zhang, Yifei Song, Claire Gardent· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.