Activation steering offers an inference-time defense for vision--language models (VLMs) by modifying intermediate representations without updating backbone parameters. However, protection on benchmark inputs may not persist across alternative expressions of the same harmful request. We investigate this gap using fixed...
Xin-Wei Zhang, Ao-Ting Hu, Hang-Cheng Liu et al.· 0 citations
As the capabilities of Vision Language Models (VLMs) continue to improve, they are increasingly targeted by jailbreak attacks. Existing defense methods face two major limitations: (1) they struggle to ensure safety without compromising the model’s utility; and (2) many defense mechanisms significantly reduce the model’...
Xiyu Zeng, Siyuan Liang, Liming Lu et al.· IEEE Transactions on Informa...· 4 citations
SafeSteer is a lightweight, inference-time steering framework that effectively defends against diverse jailbreak attacks without modifying model weights, using the innovative use of Singular Value Decomposition to construct a low-dimensional safety subspace during inference.
Xiyu Zeng, Siyuan Liang, Liming Lu et al.· arXiv.org· 0 citations
This perspective clarifies where current alignment methods genuinely benefit from game-theoretic analysis, where the framework is looser, and what challenges remain in building robust, adaptive, and verifiable AI systems.
Yaxin Cai, Zhong-Rui Zhao, Zhigang Lu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.