Preprint
Aug 2026
When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning
This work introduces Visualized Task Semantics (VTS), a controlled intervention that moves the question into the image while keeping the source problem and answer fixed, and requires no OCR or region metadata at inference.
Yongxin Wang, Ruizhe Zhou, Yueling Tang et al.
· 0 citations