Beyond Classification: Structured Supervision Aligns Visual Evidence with Medical Semantics
This work evaluates three independent forms of structured supervision: topological priors via graph self-supervision, dense pixel-level constraints via segmentation, and cross-modal semantic grounding via image-text pairs, and finds that image-text alignment achieves the most superior performance.