Skip to content
Open access

Primitive-Augmented Transformers with Event-Role Side State: Architecture Evidence, Warm-Started Modulation, and Decoupled Tool Interfaces

Jul 2026 · Machine Learning and Knowledge Extraction · Vol 8, pp. 201 · 1 citation · 18 references

TL;DR

PAT-ER, a decoder architecture with a normal token stream, an event-role register stream, and a primitive register stream, is introduced; the model is not a theorem prover and does not achieve perfect unseen tool-name copying; the contribution is a measured architecture signal and a usable, guarded interface.

Abstract

Large language models can emit fluent text while leaving intermediate semantic structure implicit. We study whether explicit event-role and logical-primitive side-state can improve a pretrained decoder without damaging its language behavior. We introduce PAT-ER, a decoder architecture with a normal token stream, an event-role register stream, and a primitive register stream. The primitive stream is motivated by the view that logical primitives answer characteristic semantic questions, such as what licenses a conclusion, what conflicts with it, or why evidence is insufficient. Across eight seeds on the same Qwen3-0.6B backbone, replacing token-pooled auxiliary heads with typed PAT-ER registers improves primitive macro-F1 by 0.209 (95% CI [0.182, 0.237]) and role-to-primitive macro-F1 by 0.091 (95% CI [0.074, 0.110]) with no language-model loss cost. A generic-register control shows that this is not merely the effect of adding latent registers: typed PAT-ER improves over generic registers by 0.116 primitive macro-F1 and 0.110 role-to-primitive macro-F1, with both confidence intervals excluding zero. A warm-started model then recovers pretrained language quality (LM loss 1.344 versus 2.555 for the frozen-backbone register model) while retaining most side-state behavior. Finally, a decoupled interface mode produces robust schema-grounded function calls on 242 held-out prompts (Hermes parse 0.952, exact arguments 0.981, JSON validity 1.000, IDK F1 1.000) while base-mode side-state metrics remain byte-identical to the warm-start baseline. The model is not a theorem prover and does not achieve perfect unseen tool-name copying; the contribution is a measured architecture signal and a usable, guarded interface.

Read PDF

Similar papers

Jul 2026

Mergeable Model-Side Aggregation States for Long-Context Language Models

A model-side aggregation interface is introduced that maintains compact Hash-based HyperLogLog sketch states alongside a frozen language model that improves over direct full-context reasoning over chain-of-thought reasoning.

Da-Chuan Song, Ju Yin, Ze-Chen Hu et al. · 1 citation
Preprint Aug 2026

Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State Injection

This work introduces PRECOG (Pre-Computed Context Injection), a retrieval mechanism that exploits a property unique to SSMs: the fixed-size, position-agnostic recurrent hidden state is a complete summary of everything the model has read.

Anusha Madan Gopal, Aras Pirbadian, Kristofor D. Carlson et al. · 0 citations
#machine learning Preprint Aug 2026

Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors

This work introduces a general steering technique called Semantic Overlays: small learned adapters applied at chosen prefill positions to a frozen model's residual stream that defends against the broad class of prompt injections that add instructions in untrusted context.

Joshua Penman · 0 citations
#machine learning Preprint Sep 2026

The Token Before the Value Is the Key: How Hybrid Architectures Organize Induction Circuits

Hybrid language models can improve capability as well as efficiency, raising the question of how architectural complementarity becomes learned computation. We examine the established induction roles of Carrying predecessor information, Matching a source by content, and Copying its value. How are these position-sensitiv...

Ke Cheng, Xin Xu, Yi-Xiao Chen et al. · 0 citations
Preprint Aug 2026

Right Reset: Chunking by Prefix Removal

Right Reset (RR) is introduced and an observed-token likelihood-ratio readout is competitive in some architectures, indicating that the central contribution is the intervention: context dependence itself can provide a boundary signal when surface structure is weak.

Mike Vegeto · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.