ConceptLens: Interactive Concept Bottlenecks for Black-Box Models
Abstract
Most interpretability tools explain model decisions post-hoc but offer limited support for user action. We present ConceptLens, a model-agnostic framework that transforms explanations into an interactive interface for probing pre-trained classifiers across domains and modalities. By extracting concept vectors and fitting local approximations, the system enables users to evaluate local fidelity, perform what-if analysis, conduct counterfactual searches, and generate evidence-grounded explanations. Unlike traditional feature-level tools, ConceptLens shifts interaction to semantic concepts rather than isolated pixels or tokens. We demonstrate the system across three tasks spanning different domains and modalities: chest X-ray classification, bird species recognition, and toxic comment detection. The framework enables users to inspect black-box decisions and also to actively refine the underlying concept probes through interactive feedback. Code available on GitHub.