Skip to content

Author

Yehui Tang

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

RoboICL: Embodied In-Context Learning with GPT-6 Astra

General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \emph{RoboICL}, an in-context robot-control framework that nar...

Fang-Cheng Liu, Ye-Qing Shen, An-Da Cheng et al. · 0 citations
Jul 2026

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification

Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, have emerged as a compelling paradigm for reliable and interpretable multimodal reasoning. However, recent studies have revealed that such models often use tools unfaith...

Haoqing Wang, Xing-Run Xing, Wei Xia et al. · 1 citation
Preprint Aug 2026

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making

Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images. Many such tasks, however, require more than rec- ognizing what an image contains: a model must use visual evidence to make a complete decision whose parts jointly satisfy global c...

Ning-Xin Pan, Han-Yu Li, Ye-Hui Tang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.