Prototype-Aware Ambiguity Mining for Tail-Sensitive Visual Recognition
Abstract
Long-tailed image datasets often produce classifiers that appear reliable under average accuracy but remain fragile on rare visual categories. Existing reweighting and focal-style objectives reduce this bias from label counts or prediction loss, yet they do not explicitly measure whether a minority-class embedding is drifting toward a visually similar majority class. This paper presents Prototype-Guided Hard Example Mining (PGHEM), a lightweight tail-sensitive training framework that maintains a dynamic multi-prototype bank for each class. PGHEM compares the closest target prototype with the strongest competing prototype, estimates a sample-level ambiguity score, and increases the loss contribution of examples located near competing class centers. A small separation term further discourages persistent prototype-level boundary overlap. On Fashion-MNIST-LT with imbalance factor 50, the two-prototype PGHEM obtains the highest mean tail-class accuracy at 91.97%, compared with 83.54% for cross-entropy, 91.11% for class-balanced cross-entropy, 91.80% for focal loss, and 90.66% for the K = 1 special case. A paired test against class-balanced cross-entropy does not show a significant tail-accuracy difference under the current five-seed variance. A detailed empirical analysis examines prototype momentum, ambiguity sharpness, separation strength, training duration, and prototype capacity.