Rewired or Gated? How Instruction Tuning Shapes Knowledge-Conflict Circuits in LLMs
In language models, the choice between believing the prompt and believing the weights is made by a handful of identifiable attention heads. Instruction tuning changes how models behave under conflict, but whether it rewires the underlying circuit or merely gates/reweights already present components, remains unknown. We...