Skip to content
Preprint

ITL: Interpretable Document Alignment with Structured Reference Frameworks

Aug 2026 · 0 citations · 22 references
Computer Science

TL;DR

Intelligent Target Locator (ITL), a domain-agnostic and language-portable methodology that estimates the affinity between the textual units of a target document and the concepts defined in a structured reference document, is presented.

Abstract

Measuring alignment between documents and structured reference frameworks requires identifying conceptual evidence distributed throughout the text and reporting it through measures that are quantitative, interpretable, and traceable. Many commonly used retrieval and classification approaches return either pairwise similarity scores or one or more class labels, whereas fewer methods provide concept-level scores that are directly traceable to the terminological evidence supporting them. We present \emph{Intelligent Target Locator} (ITL), a domain-agnostic and language-portable methodology that estimates the affinity between the textual units of a target document and the concepts defined in a \emph{Structured Reference Document} ($SRD$). From the $SRD$, ITL induces concept-specific terminological profiles built from independent terms, bigrams, trigrams, and co-occurrences. Each term is assigned an importance weight that combines concept membership, term-type specificity and inter-concept discriminability. The output is a textual-unit--concept affinity matrix that can be aggregated at different levels of granularity. We conduct an internal consistency assessment using the 17 Sustainable Development Goals (SDGs), evaluating each official goal statement against the $SRD$ induced from the same set of descriptors. Every statement reached its highest affinity with the corresponding concept, and the mean affinity across the remaining concepts stayed marginal relative to the mean reference affinity. This separation indicates that ITL distinguishes the conceptual profiles of the framework. ITL thus offers a general basis for quantifying document alignment with structured frameworks while keeping each result traceable to the terminological evidence that supports it.

View source

Similar papers

EviMap: Evidence-Grounded Hierarchical Topic Maps for Exploring Unlabeled Corpora

Research teams and organizations often explore unfamiliar free-text collections, from survey comments and reviews to reports and domain documents, before labels, queries or coding schemes exist. At this stage, the first thematic map shapes what users notice, prioritize and carry into downstream analysis, so it should b...

Zhiyin Tan, Changxu Duan · 0 citations
Book Open access Sep 2026

ReSB²: Retrieving Similar Brazilian State Bills

Legislative knowledge evolves as an intricate hypertext in which documents are interconnected through complex, often implicit relationships. In this paper, we introduce ReSB2, a framework for retrieving and linking similar legislative bills that supports human–machine collaboration and helps reduce redundancy in the la...

Lucas G. L. Costa, Átila Souza, Elves Rodrigues et al. · 0 citations
Conference Jul 2026

Metadata Matters: A Hybrid Retrieval Framework for Structured Financial Document Analysis

Large Language Models are deployed in financial applications such as research synthesis and risk analysis, yet their effectiveness is constrained by the limitations of conventional retrieval methods. Existing approaches rely primarily on semantic similarity or token-level matching, which fails in structured domains lik...

S. Jambula, Srihari Kumar Pendyala, Rajesh Kumar Butteddi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

Scientific papers require models to reason jointly over text, equations, figures, tables, code, and datasets while preserving the provenance of supporting evidence. Existing benchmarks typically evaluate these capabilities in isolation, leaving unclear whether multimodal models can support realistic scientific-reading...

Shenxi Wu, Yu-Hong Liu, Hao-Song Zhang et al. · 0 citations
Jul 2026

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost, and is the first to score value accuracy, record completeness at scale, grounding, and measured cost together.

Boyang Zhang, Adrian Lyjak, Elizabeth Stewart et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.