Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes...
Rafi Ibn Sultan, Xiang-Yu Zhou, Mohammad O. S. Chowdhury et al.· 0 citations
BACKGROUND
Accurate segmentation of the left anterior descending (LAD) artery in 3D free-breathing, non-contrast CT is critical for cardiac dose sparing in thoracic radiotherapy. The task is inherently difficult because the LAD is extremely small, exhibits poor soft-tissue contrast, and varies substantially across pati...
Rafi Ibn Sultan, Chengyin Li, Yiannos Demetriou et al.· Medical Physics (Lancaster)· 1 citation
BACKGROUND
Manual organ-at-risk (OAR) delineation takes 20-40 min per case, a major bottleneck within the 50-90 min treatment window of abdominal MR-guided adaptive radiotherapy (MRgRT). Most deep learning systems adopt single-fraction approaches that discard valuable temporal context from prior treatment fractions....
Chengyin Li, D. Rusu, Rafi Ibn Sultan et al.· Medical Physics (Lancaster)· 0 citations
Medical image segmentation is still largely treated as a vision-only problem, although clinical interpretation often relies on textual knowledge of anatomy, location, appearance, and surrounding context. Existing text-guided segmentation methods within the Vision-Language Model (VLM) paradigm often use language only as...
Rafi Ibn Sultan, Hui Zhu, Chengyin Li et al.· 0 citations
The results show that PURGE consistently reduces hallucinations and spurious-correlation-driven errors while maintaining or improving overall performance in most evaluated settings, providing both a reusable evaluation protocol and an effective mitigation framework for more reliable LVLMs.
Aditi Sarker, Nazreen Shah, Rafi Ibn Sultan et al.· 0 citations
The framework, INFUSE, first stabilizes visual and textual representations around perturbation-averaged and ground-truth anchors, then aligns the stabilized representations across modalities with bidirectional contrastive objectives.
Aditi Sarker, Rafi Ibn Sultan, Hui Zhu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.