World models have emerged as powerful technology for facilitating policy learning in robotics by providing predictive representations of environmental dynamics. A crucial component of such models is the internal state representation, which serves as a bridge between observation, decision-making, and future state predic...
Li-Xuan Zhang, Meina Kan, Shiguang Shan et al.· IEEE Robotics and Automation...· 0 citations
Advances in multimodal understanding, reasoning, and tool use enable agents to tackle increasingly complex visual reasoning tasks. By distilling past execution experience into reusable skills, agents can transfer lessons from both successes and failures into future reasoning, reducing repeated errors and improving capa...
Bei Yan, Yue-Cong Min, Jie Zhang et al.· 0 citations
Object hallucination remains a major challenge for large vision-language models. While off-policy preference optimization proves to be an effective solution, on-policy reinforcement learning provides a more promising direction as it directly targets a model's current failure modes. However, we find that without fine-gr...
Xing-Ming Long, Jie Zhang, Yue-Cong Min et al.· 0 citations
Backdoor as Probe is proposed, a test-time adversarial defense for CLIP that improves average robust accuracy from 1.0\% to 52.3\% while retaining clean accuracy, achieving performance comparable to state-of-the-art methods with up to a \(5.7\times\) inference speedup.
Zhong-Qi Wang, Jie Zhang, Sen Nie et al.· 0 citations
Adversarial attacks have long posed a fundamental threat to machine learning systems. As multimodal large language models (MLLMs) rapidly evolve and become widely deployed, assessing their vulnerability to such attacks is essential for their safe use. In this work, we investigate whether a single adversarial image can...
Sen Nie, Jie Zhang, Zhong Ling Wang et al.· 0 citations