Skip to content

5W1H+Which: Context-Valid Semantic Indexing with Progressive Ontology Binding

Sep 2026 · 0 citations · 6 references
Computer Science

TL;DR

This work proposes 5W1H+Which, a semantic indexing design that separates content extraction from ontology binding, and distinguishes business valid time, system knowledge time, and operational traces, and uses dependency records to support binding revalidation and the maintenance of derived conclusions.

Abstract

Transforming raw data into queryable knowledge requires both early extraction of reusable information and explicit types, relations, and applicability conditions for particular tasks. If indexing selects content too early around a single business schema, later tasks may be unable to use information that was omitted. If the index retains only open-ended text, however, rule-based reasoning lacks checkable premises. We propose 5W1H+Which, a semantic indexing design that separates content extraction from ontology binding. The 5W1H questions organize source-grounded content units; Which points to versioned ontology elements and records mapping relations, scope, and validation status. Time, location, system environment, and participant roles are not merely retrieval labels: together, they constrain the contexts in which facts, bindings, and rules apply. Unbound content remains searchable, while bound content enters a formal reasoning path only after premise checks. The method further distinguishes business valid time, system knowledge time, and operational traces, and uses dependency records to support binding revalidation and the maintenance of derived conclusions. A worked example of migration from an on-premises server to a cloud environment illustrates the different treatment of world-state changes, ontology-version changes, and changes in rule applicability. We formulate three groups of falsifiable hypotheses concerning cross-task evidence coverage, control of contextual misuse, and incremental update cost. The planned evaluation includes a strong typed fact-graph baseline with the same evidence, temporal information, and budget, to test whether benefits arise from 5W1H organization, deferred binding, or additional information and engineering effort. The contribution is a testable indexing mechanism, not a claim to a new universal ontology or a demonstrated performance advantage.

View source

Similar papers

Preprint Aug 2026

GrOIL: Graph-Grounded Domain Ontology Induction with Constrained LLM Mediation

A seven-stage graph-grounded pipeline that converts domain documents into a complete, auditable Web Ontology Language (OWL) Terminological Box (TBox) without any unconstrained generation step is presented, demonstrating that the pipeline produces stable, reusable domain representations from large document corpora.

Maruf Ahmed Mridul, A. Talukder, O. Seneviratne · 0 citations
#artificial intelligence Preprint Sep 2026

ArticleMiner: Ontology-Guided Knowledge Graph Construction from Scientific Publications

This work builds four modules (for drug-discovery chemistry, materials science, machine learning, and mineral geochemistry) in the ArticleMiner framework, and evaluates them on 163 papers, including a new geochemistry benchmark with expert-curated ground truth.

Md Abrar Jahin, Craig A. Knoblock, Jay Pujara · 0 citations
Preprint Aug 2026

Toward Effective and Reliable LLM Agents via Dynamic Ontology

OaK is presented, an ontology-as-a-kernel framework that dynamically constructs and refines task-oriented ontologies for LLM agents and shows that OaK improves standard LLM agents, strengthens evidence grounding, and boosts the reliability of multi-step reasoning.

Xiaohui Zhang, Ze-Qun Sun, Cheng Yang et al. · 0 citations
Preprint Aug 2026

Curate Before You Connect: Identity and Ontology Tagging in a Production Knowledge Graph

This paper describes the ingestion and ontology-tagging layer that turns a validated extraction stream into a knowledge graph of 537,157 entities and 2,198,567 relationships drawn from 98,795 government documents, and describes a record-identity ladder that decides sameness from identifier columns, name columns, displa...

Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.