Skip to content
Preprint

CBX-Bench: A Human-Aligned MLLM Council for Benchmarking Concept Bottleneck Model Explanations

Aug 2026 · 0 citations · 24 references
Computer Science

TL;DR

A multimodal large language model (MLLM) council is developed that, given an image and its CBM explanation, produces an explanation quality score, and CBX-Bench, a public benchmark and leaderboard, provides a human-aligned, scalable evaluation of CBM explanations beyond accuracy and isolated qualitative examples.

Abstract

Concept Bottleneck Models (CBMs) are designed to make visual classification interpretable by expressing predictions through human-understandable concepts. Although interpretability is the central motivation for CBMs, they are still largely evaluated as predictive models by downstream classification accuracy, supplemented by isolated qualitative examples. This highlights a pressing need for quantitative measures, a challenge complicated by the infeasibility of ground-truth concept annotation at scale and the open nature of concept lists due to a lack of consensus. To fill this gap, we develop a multimodal large language model (MLLM) council that, given an image and its CBM explanation, produces an explanation quality score. To ground and validate the council, we first conduct a human study to establish a ground-truth reference for CBM explanation quality: for an image, annotators compare explanations from two of LF-CBM, VLG-CBM, and CBM-Suite and choose the more useful one, or mark them as equally good or equally bad, yielding 2700 judgments over 900 image-comparison items on CUB-200, ImageNet-100, and Places365. Against this human reference, our five-model council, consisting of open-weight MLLMs, recovers over 70% of strict human preference rankings, rising to 83% on items where human annotators unanimously agree. Building on this validated council, we introduce CBX-Bench, a public benchmark and leaderboard: authors of new CBMs can submit their model's explanations, and CBX-Bench scores them with the council and maintains dataset-level rankings of explanation quality. CBX-Bench thus provides a human-aligned, scalable evaluation of CBM explanations beyond accuracy and isolated qualitative examples. The benchmark is available at https://github.com/meric-karadag/cbx-bench.

View source

Similar papers

Open access Sep 2026

Concept-Bottleneck Models with Expert-Discovered Versus Automatically-Discovered Concepts: A Comparative Study on Fine-Grained Fungal Classification

Interpretable-by-design architectures, such as Concept-Bottleneck Models (CBMs), are essential for trustworthy AI under regulations like the EU AI Act. While classical CBMs rely on costly expert labels, recent automated methods use vision–language models for concept discovery. However, whether automation preserves inte...

Vincenzo Bevilacqua, A. Di Marino, A. Ciaramella et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

This work introduces DocHop, a benchmark for integrated chart--context reasoning in document-style images and constructs DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, to enable systematic evaluation.

Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al. · 1 citation
#machine learning Preprint Sep 2026

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and expli...

Himil Vasava, Ming-Zhou Jiang · 0 citations
Conference Open access Sep 2025

DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values

DiverValue-Bench is introduced, a population-aware benchmark for evaluating multi-dimensional value alignment across 74 countries/regions and it is shown that lightweight preference-based fine-tuning with Low-Rank Adaptation and Direct Preference Optimization substantially improves in-domain value alignment while yield...

Yao Liang, Dongcheng Zhao, Fei-Fei Zhao et al. · 0 citations
#machine learning Preprint Sep 2026

Evaluation of Contextual Understanding in Large Language Models

A novel knowledge graph-based evaluation framework is proposed introducing Semantic Structural Similarity for KGs (S3KG), a hybrid similarity measure integrating structural and semantic similarity into a continuous evaluation score, alongside a diagnostic framework for categorizing reasoning errors.

Subavarshana Arumugam, Mamta Nallaretnam, K. Wickramasinghe et al. · 0 citations
#natural language process... Preprint Aug 2026

XQDT: eXplainable and Quantitative Data-Text Alignment Metric with Feedback Signals

This work presents an end-to-end explainable evaluation metric that fine-tunes a language model to identify omitted, extra, incorrect, and correct data units in a data-text pair, providing both fine-grained diagnostic feedback and an interpretable measure of alignment quality.

Kun Efimov-Zhang, Yifei Song, Claire Gardent · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.