Learning to Look Again: Loss-Gap Supervision for Free-form Crop Routing in Vision-Language Models
GapSight is proposed, a framework for learning visual re-reading: a VLM first takes a global glance, then selectively returns to a free-form region when the question calls for local evidence, and Mechanism analyses show that the router rescues concrete wrong answers, adapts its action rate by task, and forms a favorable token-performance profile.