Skip to content
Preprint

A Heterogeneous Mixture of Experts Framework for Interpretable Machine Learning

Aug 2026 · 0 citations · 38 references
Mathematics Computer Science

TL;DR

The proposed approach combines interpretability, adaptive inductive bias selection, and probabilistic coherence within a unified mixture-of-experts framework to achieve interpretable expert assignments while achieving predictive performance competitive with homogeneous MoDT and Random Forests.

Abstract

Mixture-of-Experts (MoE) models provide a flexible framework for partitioning complex prediction problems into simpler local learning tasks through an input-dependent gating mechanism. Existing interpretable MoE approaches, such as Mixture of Decision Trees (MoDT), achieve transparency by employing homogeneous decision-tree experts, but this restricts the model to a single inductive bias across all regions of the feature space. We extend the MoDT framework by introducing heterogeneous expert families comprising decision trees, linear support vector machines, and quadratic discriminant analysis under a common probabilistic gating mechanism. To ensure coherent likelihood-based inference, non-probabilistic experts are calibrated to produce conditional class probabilities, allowing parameter estimation within the generalized Expectation-Maximization framework of MoDT. We further establish theoretical monotone ascent guarantees for the proposed heterogeneous gating updates, providing a justification for the optimization procedure. Experiments on a diverse collection of synthetic and real-world benchmark datasets demonstrate that the proposed framework adaptively specializes experts according to local data geometry, yielding interpretable expert assignments while achieving predictive performance competitive with homogeneous MoDT and Random Forests. The proposed approach combines interpretability, adaptive inductive bias selection, and probabilistic coherence within a unified mixture-of-experts framework.

View source

Similar papers

#machine learning Preprint Sep 2026

Towards a Statistical Understanding of Mixture-of-Experts

Mixture-of-experts (MoE) architectures increase model capacity by combining a collection of expert predictors through input-dependent routing, while often activating only a small subset of experts for each input. Despite their growing importance in modern large-scale models, the statistical roles of their design choice...

Si-Yuan He, Bo-Kai Yang, Jie Hu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts

Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling model capacity while preserving efficient inference in large foundation models. However, most MoE models use a fixed top-$k$ expert selection policy, assigning the same expert budget to every token even when fewer experts may be suff...

Dohyeon Kim, Bedionita Soro, Sung Ju Hwang · 0 citations
#machine learning Preprint Aug 2026

Structure Aware Neural Architecture Search for Mixture of Experts

An architecture search framework that makes the alignment between experts and the structure of the data an explicit search variable and ensures that the assignment of data clusters to experts is optimised jointly with the per-expert architectures is proposed.

Petr Babkin, O. Bakhteev · 0 citations
#machine learning Preprint Aug 2026

Enhancing Bayesian Optimization and Active Learning Through Kernel Diversity

A unified framework, KENDO (Kernel ENsemble Disagreement-aware Operator), is proposed that integrates Ensemble Gaussian Processes (EGP) with disagreement-aware acquisition strategies and extends the approach to multi-objective optimization via random scalarization that preserves the single-optimizer conditioning struct...

Heng Zhang, Hao-Tian Xiang, Konstantinos D. Polyzos et al. · 1 citation
Preprint Aug 2026

Semiparametric robust mixture of experts based on nonparametric maximum likelihood

The mixture of experts (MoE) model provides a flexible approach for modeling heterogeneous regression relationships by allowing covariate-dependent mixing through a gating network, but most existing MoE models rely on parametric assumptions for expert error distributions, typically Gaussian, which can lead to inefficie...

Sangkon Oh, V. H. Lachos, Byungtae Seo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.