Skip to content

Machine-Interpretable Information: Compiling Documents into Searchable and Readable Protocol States

Sep 2026 · 0 citations · 67 references
Computer Science

TL;DR

This work introduces Machine-Interpretable Information (MII), the first agent-to-agent (A2A) document-to-state protocol, and demonstrates strong cross-model interoperability across heterogeneous LLMs -- despite the Writer using a legacy GPT-2 vocabulary, forcing genuine semantic translation rather than token-level memorization.

Abstract

Long-context language models interface with external knowledge through raw natural language. In retrieval-augmented systems, this creates a persistent index-payload schism: dense vectors enable searchable routing, but models must re-ingest lengthy text payloads for reasoning at O(N^2) attention cost. Existing compression methods further produce private states tied to specific architectures. We introduce Machine-Interpretable Information (MII), the first agent-to-agent (A2A) document-to-state protocol. A dual-timescale state-space Writer compiles documents into a canonical, fixed-bandwidth state (56 tokens), and a lightweight Translator maps it into any frozen Reader's embedding space, reducing query-time cost to O(K). The resulting .mii artifact unifies Retrieval (searchable geometry), Reasoning (global memory), and Reconstruction (grounded details) in a single transferable medium. We demonstrate strong cross-model interoperability across heterogeneous LLMs (e.g., Llama, Qwen, Mistral) -- despite the Writer using a legacy GPT-2 vocabulary, forcing genuine semantic translation rather than token-level memorization. Mechanistic probes reveal modular latent structure: entity representations can be causally traced and zero-shot transplanted between unrelated document states while remaining decodable. To address lexical reconstruction under fixed bandwidth, we propose Residual-MII, a cache hierarchy combining compiled global memory with sparse local evidence. On HotpotQA (7,405 queries), Residual-MII exceeds full-context Exact Match at approximately 7% of the attention FLOPs, suggesting a paradigm shift toward compiled, transferable neural document formats.

View source

Similar papers

Preprint Aug 2026

Code as Representation: A Compilable Parsing Paradigm for Academic Documents

Compilable Academic Document Parsing (CADP) is proposed, a paradigm that reconstructs a full page as contextual \LaTeX{} plus executable Python, so that structure-preserving elements and executable chart representations can be reconstructed, recompiled, and directly verified against the source page.

Rihui Jin, Jun Wang, Chen Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Retrieval-Augmented Generation for Scientific Code Understanding

The results indicate that front-loading code understanding into a reusable, codebase-specialised vector store enables small local models to deliver grounded and repository-specific answers, making the agent well suited as a privacy-preserving development tool for in-house scientific codebases.

Aaron Nobile, Andreas Adelmann, Mohsen Sadr · 0 citations
#natural language process... Preprint Sep 2026

Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning

The Context Compilation Architecture (CCA), whose central novelty is a typed intermediate representation (IR) with fixed slots (rules.{must_do, must_not, conditional}, output_spec, available_tools, data_profile) into which any prose context is compiled once; executable verifiers and a violation-gated correction loop fo...

Jin-Hu Qi, Min-Da Hu, Wen-Tao Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis

Datalog underpins reasoning tasks such as program analysis, but its programs are hard to write. Existing synthesizers automate this task but require users to state their intent as input-output examples. Large language models (LLMs) suggest a more natural route, text-to-Datalog synthesis from a natural-language question...

Yuan Li, Han-Yun Jiang, Guo-Wei Tian et al. · 0 citations
Review Open access Sep 2026

CITAE: A Passage-Verifiable Retrieval-Augmented Generation Platform for Researchers and Thesis Writers

CITAE is an open-source web application that integrates the complete scientific-literature workflow—discovery, read- ing, verification, organisation and citation—into a single system, and makes AI grounding visible to the user rather than hidden inside the model.

Richar Andre Vilca-Solorzano, Dina Maribel Yana-Yucra, Fred Torres-Cruz et al. · 0 citations
Book Open access Aug 2026

Docling: Converting Complex Documents into AI-Ready Structured Representations

By bridging the gap between visually complex documents and machine-readable knowledge, Docling provides a foundation for reliable document understanding in next-generation AI systems.

P. Staar · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.