Skip to content

A Mechanistic Lens on Semantic Conflicts: Using Activation Patching to Understand LLM Behavior

Jul 2026 · arXiv.org · Vol abs/2607.05587 · 2 citations · 55 references
Computer Science

TL;DR

The findings show that semantic conflicts affect program comprehension and downstream tasks, with relevant information concentrated in a small number of causally active residual-stream states, and demonstrate a framework for mechanistically analyzing how LLMs integrate code-related information under controlled semantic variations.

Abstract

Large language models (LLMs) are increasingly used in software-engineering tasks processing executable code and non-executable semantic cues such as comments or identifiers. These two sources can conflict when semantic cues suggest different program behavior than the code itself. It remains unclear how such semantic conflicts affect LLM behavior and which source dominates their outputs. We present the first controlled, mechanistic study of LLM behavior under semantic conflicts. To this end, we construct 45 Python snippet triplets that isolate conflicts by varying either semantic cues or implementation while keeping token-aligned pairs for causal intervention. We evaluate four open-weight LLMs on two tasks (output prediction and unit-test generation) using behavioral performance measures and residual-stream activation patching to identify token-layer states that causally contribute to behavioral differences between aligned and conflicting inputs. Our results show that semantic conflicts significantly reduce execution-grounded correctness in both tasks and that all tested LLMs often follow misleading semantic cues. Residual-stream activation patching reveals a consistent pattern for final-output prediction: The changed cue/code region and a small set of intermediate tokens carry most of the recoverable causal signal before aggregation near the output readout. For unit-test generation, this pattern extends beyond the prompt, showing that conflict-related information is recoverable at generated sites before producing expected values. Overall, our findings show that semantic conflicts affect program comprehension and downstream tasks, with relevant information concentrated in a small number of causally active residual-stream states, and demonstrate a framework for mechanistically analyzing how LLMs integrate code-related information under controlled semantic variations.

View source

Similar papers

Preprint Aug 2026

Why Do LLMs Fail at OCL Generation? A Graph Reasoning Perspective

The underlying causes of LLM failures in OCL generation are investigated, framing the task as a graph reasoning problem over UML class diagrams and finding that OCL generation performance significantly degrades with increasing navigation depth and structural complexity.

Hamza Attarwala, Moataz Chouchen, Mohammad Hamdaqa et al. · 0 citations
Preprint Aug 2026

Evaluating Language Models on Cross-Language Code Functional Equivalence

This work investigates whether LLMs can accurately judge functional equivalence across different programming languages in human-written code, a setting that requires deeper reasoning beyond superficial similarity, and identifies a difficulty-dependent breakdown in equivalence judgment.

Hui Sun, Anderson G. Uchôa, Rohit Gheyi et al. · 2 citations
#artificial intelligence Book Open access Sep 2026

Predicting Program Exit Code with LLMs and Programming Language Semantics

This work evaluates open-source coding LLMs under two semantic formalisms and two semantic shifts across Human-Written, LLM-Translated, and Fuzzer-Generated program splits, showing that LLMs lean on pre-training priors rather than systematically applying the given rules.

Lara Marinov, Aditya Thimmaiah, Jayanth Srinivasa et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

SWE-Flux is introduced, a repository-level benchmark for dynamic execution reasoning containing 480 execution-grounded instances across 12 real Python repositories, with gold answers automatically harvested from instrumented test executions rather than written manually or judged by LLMs.

Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband et al. · 0 citations
#machine learning Preprint Sep 2026

On the Lexical Superstition of Large Language Models for Code Comprehension: Re-evaluation on Code of Low Lexical Quality

Face/Off is introduced, a semantics-preserving identifier-renaming framework, and progressive naming conditions across multiple models and code-comprehension tasks are evaluated, revealing a systematic vulnerability in how current LLMs balance lexical cues against program structure.

Xin-Peng Shen, San-Zhuo Xi, Ya-Li Du et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.