Beyond the Tumor Signature: Investigating Global Mammographic Context in Machine-Learning–Assisted Breast Imaging
Abstract
Artificial intelligence (AI) is increasingly studied for mammographic image analysis, but model performance depends on preprocessing, dataset design, and evaluation choices. This study tested whether focusing a traditional machine-learning pipeline on an annotated lesion improves classification of benign versus malignant abnormalities. A Kaggle-distributed copy of the mini-MIAS mammography dataset was used, with dataset identity and annotation definitions checked against the MIAS documentation [1]. Abnormal mammograms with valid lesion coordinates were analyzed using two strategies: the whole mammogram and a lesion-centered region of interest (ROI). Both used identical Histogram of Oriented Gradients (HOG) features [2] and class-weighted logistic regression. Performance was evaluated with 10 repeats of 5-fold stratified, patient-grouped cross-validation. Whole-mammogram processing produced mean ROC-AUC 0.571 (95% CI 0.540–0.602), compared with 0.532 (95% CI 0.496–0.567) for ROI processing. The paired ROI-minus-whole difference was −0.039 (95% CI −0.080–0.001; paired t-test p=0.055). Sensitivity was 0.537 for whole images and 0.498 for ROI images. Lesion-focused cropping therefore did not improve this HOG-plus-logistic-regression