Skip to content

Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure

Sep 2026 · 0 citations · 15 references
Computer Science

TL;DR

A hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate channels, hold the token sequence fixed, intervene on the paragraph coordinate $p_1$, and measure cross-paragraph attention with a token-distance-exact estimator is used.

Abstract

Standard positional encodings treat position as a one-dimensional reading-order coordinate, but reading order alone does not determine hierarchical textual structure. We use a hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate channels, hold the token sequence fixed, intervene on the paragraph coordinate $p_1$, and measure cross-paragraph attention with a token-distance-exact estimator. Attention is compressed in every corpus, but compression alone is not diagnostic of true structure: an architecturally identical channel with density-matched random labels is compressed too. What distinguishes real structure is the depth of compression: the real-versus-random depth gap is resolvable in two of three corpora and not in the third, and real-structure depth varies more strongly across corpora than the control's. Comparing eight corpus-only quantities across three constructs (lexical persistence, paragraph length, embedding-based coherence), none fully reproduces the cross-corpus ordering of depth, though embedding-based coherence comes closest. At the paragraph level that relation transfers as a common slope, but with the opposite sign to the corpus-level ranking: controlling for length, diversity, and position, paragraphs with more similar neighbors compress less deeply, with no detectable slope difference between any pair of corpora, while length, lexical-diversity, and position effects remain corpus-specific. Compression depth, not its location, is the reproducible signature of genuine paragraph structure in our setting.

View source

Similar papers

#machine learning Preprint Sep 2026

Content-Based Addressing for Long Context

It is proved that this construction preserves local RoPE exactly, leaves the attention comparison between two fixed tokens unchanged when other units are inserted or reordered, and does not create new relative rotations merely because more units are added.

M. Godavarti · 0 citations
#machine learning Preprint Sep 2026

The Residual Stream's Effective Depth

We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one number. Across sixteen decoder-only language models, $\Deff$...

Barak Gahtan, Ido Galil, A. M. Bronstein · 0 citations
#natural language process... Preprint Aug 2026

All You Need Is Non-Commutative Words

It is shown that the noncommutativity of matrix product captures word order without positional encodings (PEs) and yields several capabilities, including antisymmetric self-attention with no query, key, or value projections, and parallel composition of variable-length text chunks at a reduced attention cost.

Carla Quispe Flores, Stanley Salvatierra, Renan Cabrera · 0 citations
#artificial intelligence Review Sep 2026

MatLoom: Layered Text-to-Material Generation in a Compact Program Space

Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. Each program composes alpha-masked layers whose shared spatial expressions define coverage and physic...

Anson Y. Lam, Shu-Qing Li, M. Lyu · 0 citations
Preprint Aug 2026

TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents

TongGuOCR is proposed, a layout-aware and token-augmented multimodal large language model (MLLM) for OCR of Chinese historical documents that outperforms representative traditional task-specific OCR models, general-purpose MLLMs, and OCR-oriented MLLMs.

Zhongheng Zhou, Yi Sun, Huiguo He et al. · 0 citations
Preprint Aug 2026

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

PaDoc is proposed, a layout-grounded parser that treats the predicted layout as a branching structure over a shared page representation that is the fastest end-to-end parser at five concurrency levels and is the fastest end-to-end parser at five concurrency levels.

Hao Yu, Jia-Bo Zhan, Kang Liu et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.