Skip to content

A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models

Jul 2026 · arXiv.org · Vol abs/2607.27910 · 0 citations · 15 references
Computer Science

TL;DR

The results show that the two recipes recover partially overlapping geometry and that direction based defences should be calibrated separately for each language decoder family.

Abstract

Inference time defences against vision language model jailbreaks often subtract a calibrated direction from the residual stream at a chosen decoder layer. We compare five defence candidates across 15 model and layer cells from four architectural families under a magnitude controlled protocol that matches the intervention size for each prompt and pairs every direction with a random control of the same norm. The candidates are the mean image conditioning shift, a CMRM style refusal direction, a ShiftDC style attack specific residual, a prompt instruction to ignore the image, and a random control. No single candidate dominates on both refusal recovery and utility preservation. The image conditioning shift leads on LLaVA 1.5 and Pixtral 12B and is the only candidate whose utility loss remains at the measurement noise floor in every family. The prompt instruction leads on Qwen2.5 VL, while the attack specific residual leads on Qwen2 VL 2B. The image conditioning direction is direction specific in 13 of 15 cells, but strongly architecture specific and nontransferable across the only dimension compatible pair, LLaVA 1.5 13B and Pixtral 12B. We also connect text only and multimodal refusal geometry. The CMRM direction has positive cosine alignment with the image conditioning shift in all 15 cells, with mean 0.35, range 0.17 to 0.65, 15 to 25 times the random vector null, and a sign test p value of about 3e-5. These results show that the two recipes recover partially overlapping geometry and that direction based defences should be calibrated separately for each language decoder family.

View source

Similar papers

Preprint Aug 2026

ReACT-CLIP: Response-Aware Test-Time Defense for Vision--Language Models

This work introduces ReACT-CLIP, a response-conditioned test-time defense that separately determines how strongly each input should be corrected and whether defensive intervention is necessary, and quantifies this variation using a prediction-instability score computed by Jensen--Shannon divergence and combines it with...

H. Malik, Toluwani Aremu, Samuele Poppi et al. · 0 citations
Preprint Aug 2026

Reasoning Errors Have a Region and a Direction in the Residual-Stream Trajectory of LLMs

It is suggested that reasoning validity is better read from state-conditioned motion than from either static states or decontextualized trajectories alone, andlations show that motion, region, and direction provide complementary signals.

Hamed Damirchi, I. M. De La Jara, D. Ranasinghe et al. · 0 citations
Review Aug 2026

Which Source Wins? Task-Dependent Reliance in Vision-Language Models

Modality reliance in VLMs is not fixed, but varies across tasks, evidence structures, models, and evaluation settings, which shows that modality reliance in VLMs is not fixed, but varies across tasks, evidence structures, and evaluation settings.

Rodela Ghosh, Aviral Gupta, Guang-Jing Wang · 0 citations
#machine learning Preprint Sep 2026

Locating and Steering Refusal Beyond Attention

Where inside a language model does refusal live, and does that place change when the architecture does? In a transformer, refusal is governed by a single direction in the residual stream, a finding that safety and interpretability tooling now depend on. State-space models (SSMs) route information through a recurrent up...

Preethi Carmel Bosco, Gopalakrishnan Srinivasan · 0 citations
#artificial intelligence Preprint Sep 2026

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

Directional ablation removes an aligned language model's ability to refuse by projecting a single"refusal direction"out of the weights that write the residual stream. It needs no gradient-based training and no optimization, only a few hundred contrastive prompts, which makes it the canonical white-box attack on open-we...

Yi Shi, Tan-Yu Chen, Kai Shen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.