D-IMM: Distributed Iterative Mistake Minimization
Abstract
As machine learning systems are increasingly deployed on large-scale data, the demand for interpretable and scalable explanation methods has become critical. Existing Explainable AI techniques, particularly for clustering, often struggle with scalability and generalization beyond small, single-node environments. This paper introduces D-IMM (Distributed Iterative Mistake Minimization), a distributed variant of the Iterative Mistake Minimization (IMM) algorithm, designed to provide global, post-hoc explanations for k-means clustering. D-IMM constructs interpretable decision trees through efficient binning, parallelized histogram-based split selection, and mistake-driven refinement, enabling it to scale effectively with data volume and cluster complexity while maintaining high fidelity to the original clustering structure. Experimental evaluations on large benchmark datasets demonstrate that D-IMM reduces explanation error, achieves significant runtime improvements, and preserves interpretability across increasing dataset sizes and cluster counts.