Skip to content
Preprint

Align, Unify, Suppress, Route: A Coherentist View of Transformer Computation

Aug 2026 · 0 citations · 67 references
Computer Science

TL;DR

Coherentist Probabilistic Compositionalism (CPC), an interpretive framework that grounds transformer computation in coherentist theories of interpretation and describes it through four operator roles, is introduced.

Abstract

Mechanistic interpretability has identified transformer circuits, but lacks a shared vocabulary for describing how their functions compose across tasks and architectures. We introduce Coherentist Probabilistic Compositionalism (CPC), an interpretive framework that grounds transformer computation in coherentist theories of interpretation and describes it through four operator roles. Alignment identifies candidate relations, unification integrates supporting information, suppression reduces incompatible alternatives, and routing carries selected information to the output. Across 15 models from five architecture families, the suppression, unification, and routing weight-space signatures correlate with held-out activation-level role measures above random baselines. Suppression is more stable across tasks than unification. Ablating alignment heads reduces downstream suppressive activity beyond a random-head control in 10 models, but similar effects on no-conflict prompts indicate a general upstream dependency, not contradiction-specific coupling. Explicit contradictions significantly shift a layerwise coherence proxy in 14 models; after removing shared residual covariance, the gap has the predicted direction in every model. Base and instruction-tuned variants preserve induction-head score structure ($r{\geq}0.98$) without a consistent shift of operator signatures towards later layers. These results support CPC as a shared vocabulary for comparing transformer mechanisms while showing that their depth and geometric expression remain architecture-specific.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

The Geometry of Logic: Stratification Induces Semantic Structure and Robust Reasoning

Transformer-based language models perform well on symbolic tasks, yet it remains unclear whether they learn generalizable rules or rely on statistical shortcuts. Mechanistic studies link algorithmic behavior to structured internal representations, motivating the hypothesis that robust reasoning benefits from separating...

Cristina V. Lopes, Yuan-Gang Li, Justin Tian Jin Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering

Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Conventional activation edits often change both. We introduce an editable semantic-value interface on frozen residual states, with a one-way semantic-to-value pathway that grounds...

Jia-Le Dai, Hong-Can Deng, Liu-Xian Ma et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Causal and Interpretable Structures in LLM Compositional Tasks

This work studies activations from ensembles of prompts that require inferring relationships between three tokens corresponding to a cyclic concept to correctly predict the next token and finds a consistent layerwise progression in how the joint dependence among the tokens is geometrically organized and causally used.

Gurbir Arora, To-Ni J. B. Liu, Jiajun Bao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LLM Alignment--Utility Asymmetry under Semantic-Preserving Transformations

Synthetic semantic-preserving transformations that are rule-based and invertible are used as a probe of alignment generalization and suggest that semantic-preserving distribution shifts can expose a recurring gap in how utility and alignment generalize in current LLMs.

Mo-Han Li, Cheng-Yu Yu, Francesco Sovrano et al. · 0 citations
#natural language process... Preprint Sep 2026

Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs

Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal decomposition, and deduction. Although these operations are explicitly distinguished in text, little is known about how they are geometrically organized in representation spaces. To this end, we investiga...

Seogyeong Jeong, Jaehui Hwang, Dongyoon Han et al. · 0 citations
Preprint Aug 2026

Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching

Large Language Models (LLMs) demonstrate remarkable problem-solving capabilities when guided by Chain-of-Thought (CoT) prompting, yet the internal mechanisms underlying these improvements remain poorly understood. In this work, we investigate where CoT-related causal effects emerge across the generated reasoning trajec...

Murat Dura, Serkan Öztürk, Selma Tekir · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.