A Hierarchical Lesion-Conditioned Cross-Attention Network for Multimodal Breast Cancer Classification
Abstract
Existing multimodal approaches for breast cancer classification rely on fixed-stage fusion, where clinical features are incorporated as static auxiliary inputs, limiting dynamic interactions between visual and clinical information. Furthermore, reported performance gains in this literature are rarely validated under leakage-safe evaluation protocols or statistically tested against baselines, leaving true diagnostic benefit difficult to verify. To address this limitation, we propose HILCANet (Hierarchical Lesion-Conditioned Cross-Attention Network), a multimodal deep learning architecture designed to reflect the sequential and conditional nature of radiologic decision-making. The model introduces a two-level hierarchical cross-attention mechanism. At the first level (Lesion-to-Global), a lesion token interacts with global patch tokens from a pretrained encoder, enriching lesion representation with full-image spatial context. At the second level (Clinical-to-Lesion+Global), a clinical token conditions the lesion representation by integrating structured radiologic metadata. The resulting token set is aggregated using a learned attention-based pooling mechanism. We evaluate the model on the CBIS-DDSM dataset under a leakage-safe protocol, including mass-only evaluation, patient-level stratified splitting, and exclusion of post-examination labels. HILCANet achieves an AUC-ROC of 0.8711 (95% CI: 0.833–0.905), with sensitivity of 0.884 and specificity of 0.701, obtained from a five-fold ensemble fixed in advance and evaluated once on a held-out test set. Patient-level evaluation yields an AUC of 0.872. The proposed model outperform internal baselines and re-implemented prior fusion architectures evaluated under the same leakage-safe protocol. Ablation studies confirm the importance of all architectural components, while multi-seed experiments demonstrate training stability. The attention-weighted pooling mechanism provides interpretable insights. These results show that the proposed hierarchical fusion framework improves classification performance while providing rigorously validated and clinically meaningful and interpretable insights for mammography.