Skip to content
Preprint

HexMIL: Hierarchical Attention MIL for Ante-Hoc Explainable Detection of AI-Manipulated CT Volumes

Aug 2026 · 0 citations · 48 references
Computer Science

TL;DR

Hierarchical EXplainable Multiple Instance Learning (Hierarchical EXplainable Multiple Instance Learning), a mask-free medical deepfake detector that simultaneously addresses both limitations using only binary volume-level supervision.

Abstract

The emergence of medical deepfakes, i.e., medical images manipulated by deep generative models, poses a significant threat to clinical workflows. However, existing detectors suffer from two critical limitations: poor generalization to unseen generative architectures for manipulation detection and lack of interpretability. In this context, we present HexMIL (Hierarchical EXplainable Multiple Instance Learning), a mask-free medical deepfake detector that simultaneously addresses both limitations using only binary volume-level supervision. HexMIL decomposes each CT volume into a two-level hierarchy of patches and slices, aggregated via independent Gated Attention modules whose weights are directly combined into a full-resolution 3D attention volume that localizes the manipulated sub-region without any pixel-level annotation. Unlike post-hoc methods such as Grad-CAM, HexMIL's attention weights constitute the exact forward computation driving the classification decision, providing ante-hoc and structurally faithful spatial attribution. We evaluate HexMIL on M3DSynth and CT-GAN datasets under a rigorous cross-generator generalization protocol, training on a single generative architecture and testing on unseen ones. HexMIL outperforms all baselines by $+9.1$ AUC and $+9.4$ F1 in out-of-domain classification, and achieves the best average IoU and Pointing Game score in localization. Project page: opontorno.github.io/hexmil.

View source

Similar papers

Open access Aug 2026

LiteRenalNet: A lightweight dual-branch network with channel–spatial co-attention for interpretable kidney CT classification

Automated classification of renal computed tomography (CT) can accelerate triage and extend reliable screening to settings where specialist expertise is scarce, yet the most accurate deep models rely on heavy general-purpose backbones, exploit only a single visual scale, and offer little insight into the basis of their...

Hui-Hui Xie, Taifeng Yao, Jin-Feng Cao · 0 citations
Preprint Sep 2026

AlphaRAD: Grounded Zero-Shot Classification in Chest Radiology via $\alpha$-Corrected Binary Cross Entropy and Factorized Latent Supervision

Vision-Language Pretrained Models (VLPMs) offer a scalable path to open-vocabulary chest radiology understanding, yet two aspects remain underexplored: how structured clinical semantics extracted from medical reports can reduce in-batch noise during contrastive learning, and how cross-modal fusion can be designed to pr...

Jianzhong You, Yuan Gao, Chris McIntosh · 0 citations
Open access Oct 2026

HierTAC-Net: a robust and explainable hierarchical attention network for multi-configuration gastrointestinal endoscopic image classification

Colorectal cancer is a leading cause of cancer death worldwide, even though colonoscopy is an effective screening tool. Automated analysis of endoscopic images can help doctors find more lesions and make faster decisions, but current systems usually handle just one task, are tested on clear images from a single dataset...

Md. Masum Mia, A. Khandakar, S. H. Ali et al. · 0 citations
Preprint Sep 2026

Beyond Classification: Structured Supervision Aligns Visual Evidence with Medical Semantics

This work evaluates three independent forms of structured supervision: topological priors via graph self-supervision, dense pixel-level constraints via segmentation, and cross-modal semantic grounding via image-text pairs, and finds that image-text alignment achieves the most superior performance.

He-Xiang Bai, Han-Yang Xu, Xiao-Xue Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Reading the Whole Heart: Latent-Attention Masked Autoencoders for Multimodal Cardiac Representation Learning

Cardiovascular diagnosis and treatment rest on integrating complementary modalities, such as electrocardiogram, echocardiography, and chest X-rays, each capturing distinct but complementary aspects of cardiac pathophysiology. Yet most medical foundation models remain modality-specific, combining modalities only for fin...

Andrea Agostini, Simon Böhi, Moritz Vandenhirtz et al. · 0 citations
Conference Aug 2026

X-AstroNet: Explainable Adaptive Token Fusion via Hybrid Convolutional-Swin Transformer Pipelines for Alzheimer's Stage Identification

The challenge of early detection of Alzheimer's Disease (AD) and Mild Cognitive Impairment (MCI) is a neuroimaging challenge that has not been solved yet, as conventional deep learning architectures are not well suited to extract fine-grained localized features from the image while preserving the long-range structural...

S. Lokesh, P. Muthukumaraswamy, S. Priyan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.