Mitigating Hallucination in Long Referring Expressions via Training-Free, Anchor-Preserved Visual Grounding
Long referring expressions create two coupled sources of hallucination in visual grounding. A detector can select an object that matches only part of the instruction, while a structured vision–language model (VLM) branch can hallucinate a target head or an attribute–object binding. We propose DeRecG, a training-free, a...