Skip to content

Author

Xilin Chen

We have 6 of 72 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Nov 2026

Improving Policy Learning via Language-Guided State Representation in World Models

World models have emerged as powerful technology for facilitating policy learning in robotics by providing predictive representations of environmental dynamics. A crucial component of such models is the internal state representation, which serves as a bridge between observation, decision-making, and future state predic...

Li-Xuan Zhang, Meina Kan, Shiguang Shan et al. · 0 citations
Preprint Sep 2026

V-Gym: Enhancing Agentic Visual Reasoning via Skill-Data Co-Evolution

Advances in multimodal understanding, reasoning, and tool use enable agents to tackle increasingly complex visual reasoning tasks. By distilling past execution experience into reusable skills, agents can transfer lessons from both successes and failures into future reasoning, reducing repeated errors and improving capa...

Bei Yan, Yue-Cong Min, Jie Zhang et al. · 0 citations
Preprint Sep 2026

Beyond Saying Less: Fine-Grained Alignment for Informative and Faithful Vision-Language Models

Object hallucination remains a major challenge for large vision-language models. While off-policy preference optimization proves to be an effective solution, on-policy reinforcement learning provides a more promising direction as it directly targets a model's current failure modes. However, we find that without fine-gr...

Xing-Ming Long, Jie Zhang, Yue-Cong Min et al. · 0 citations
#small language model Preprint Sep 2026

Backdoor as Probe: Test-Time Adversarial Defense for CLIP

Backdoor as Probe is proposed, a test-time adversarial defense for CLIP that improves average robust accuracy from 1.0\% to 52.3\% while retaining clean accuracy, achieving performance comparable to state-of-the-art methods with up to a \(5.7\times\) inference speedup.

Zhong-Qi Wang, Jie Zhang, Sen Nie et al. · 0 citations
#machine learning Preprint Sep 2026

One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs

Adversarial attacks have long posed a fundamental threat to machine learning systems. As multimodal large language models (MLLMs) rapidly evolve and become widely deployed, assessing their vulnerability to such attacks is essential for their safe use. In this work, we investigate whether a single adversarial image can...

Sen Nie, Jie Zhang, Zhong Ling Wang et al. · 0 citations
Jul 2026

Dual Adversarial Fine-tuning for Enhancing Robustness of Large Vision Language Model

While Large Vision-Language Models (LVLMs), represented by LLaVA and GPT-4V, have demonstrated remarkable capabilities, their visual inputs remain vulnerable to adversarial attacks, posing significant security risks. Existing defense methods predominantly target single-task scenarios (e.g., zero-shot classification) an...

Sibo Wang, Jie Zhang, Shiguang Shan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.