Skip to content

Navigating the MIL Trade-Off: Flexible Pooling for Whole Slide Image Classification

2025 · Neural Information Processing Systems · pp. 107479-107524 · 1 citation · 111 references
Computer Science

TL;DR

Maxsoft is introduced—a novel MIL pooling function that enables flexible control over this trade-off, allowing adaptation to specific tasks and datasets, and PerPatch augmentation is proposed—a simple yet effective technique that enhances model robustness.

Abstract

Multiple Instance Learning (MIL) is a standard weakly supervised approach for Whole Slide Image (WSI) classification, where performance hinges on both feature representation and MIL pooling strategies. Recent research has predominantly focused on Transformer-based architectures adapted for WSIs. However, we argue that this trend faces a fundamental limitation: data scarcity. In typical settings, Transformer models yield only marginal gains without access to large-scale datasets—resources that are virtually inaccessible to all but a few well-funded research labs. Motivated by this, we revisit simple, non-attention MIL with unsupervised slide features and analyze temperature-β -controlled log-sum-exp (LSE) pooling. For slides partitioned into N patches, we theoretically show that LSE has a smooth transition at a critical β crit = O (log N ) threshold, interpolating between mean-like aggregation (stable, better generalization but less sensitive) and max-like aggregation (more sensitive but looser generalization bounds). Grounded in this analysis, we introduce Maxsoft—a novel MIL pooling function that enables flexible control over this trade-off, allowing adaptation to specific tasks and datasets. To further tackle real-world deployment challenges such as specimen heterogeneity, we propose PerPatch augmentation—a simple yet effective technique that enhances model robustness

View source

Similar papers

Preprint Aug 2026

Test-Time Instance Selection for Improved Whole Slide Image Analysis

Whole Slide Image (WSI) analysis has been widely studied for cancer diagnosis. Conventionally, a gigapixel WSI is divided into small patches and processed by Multiple Instance Learning (MIL) models. However, existing MIL models typically process all patches, many of which contain redundant or non-informative tissue pat...

Quoc Anh Nguyen, Sunhong Park, Jin Tae Kwak · 0 citations

Rethinking BCE Loss for Multi-Label Image Recognition with Fine-Tuning

Class-wise Covariance Regularization is proposed, which aligns the predicted covariance structure of class confidences with the semantic correlations encoded in pretrained text embed-dings with the geometric consistency of the class space throughout fine-tuning, resulting in more stable and interpretable confidence dis...

Ao Zhou, Zhi-Wei Jiang, Zi-Feng Cheng et al. · 0 citations
Preprint Aug 2026

PatchGen: Learning Soft Intra-Image Predictive Subsets for Visual Generalization

Visual classifiers are expected to generalize under data shifts, target shifts, and their combinations, yet most existing methods focus on domain invariance while failing to address intra-image predictive sufficiency. We investigate the structural hypothesis that each image contains a sample-adaptive oracle intra-image...

Zhaorui Tan, Weimiao Yu, Xi Yang · 0 citations
Preprint Aug 2026

DeCo-MIL: Debiased Counterfactual Reasoning for Long-Tailed Whole Slide Image Analysis

Multiple instance learning (MIL) is widely used for weakly supervised whole slide image (WSI) analysis. However, under long-tailed distributions, MIL-based WSI analysis faces a nested dual long-tail: an inter-slide class long tail and an intra-slide long tail of instance-level discriminative evidence. The two long tail...

Xiao-Xiao Li, Xi-Tong Ling, Jiawen Li et al. · 0 citations
Conference Open access Sep 2026

GroupMIL: Semantic Group Based Multiple Instance Learning for Whole Slide Image Analysing

Whole Slide Image (WSI) analysis faces challenges due to gigapixel resolutions and slide-level weak supervision. Multiple Instance Learning (MIL) serves as a pivotal method for this task. However, existing MIL frameworks often fail to exploit the inherent redundancy of tissue patterns or the semantic coherence among si...

Zhao Yao, Zhen-Mi Xie, Meng-Xin Tian et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.