· Balisage Series on Markup Technologies· 0 citations
TL;DR
The format, its schema mechanism, and its packaging are introduced: economical enough for an agent to process, and provable enough for its answer to be trusted for enterprise document AI.
Abstract
DGML — Document Graph Markup Language — is a semantic XML representation of business documents. Docugami, with Inveniam, is opening the format itself to the industry as a potential standard, with a reference implementation. Where raw extraction gives you fields, and structural markup gives you shape, DGML gives you meaning: tags that describe what each element
is
in its domain — a liability cap, an effective date, a payment obligation — not how it appeared on the page. DGML's headline property is cross-document tag consistency: a stable vocabulary across every document of a given type, making a corpus queryable without per-document prompt engineering. A four-layer architecture — semantic tagging, spatial grounding via pixel bounding boxes, cryptographic fragment-level attestation, and browser-native readability — makes DGML suitable for enterprise document AI: economical enough for an agent to process, and provable enough for its answer to be trusted. This paper introduces the format, its schema mechanism, and its packaging.
This paper traces a personal and technical journey from the HTML Document Object Model of 1996 through XML, XSLT, XSD, XSD, XQuery, RDF, SHACL, and the neural systems of the present day, arriving at the holonic graph as the resolution of a tension that has persisted, largely unacknowledged, throughout the life of the m...
K. Cagle· Balisage Series on Markup Te...· 0 citations
LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost, and is the first to score value accuracy, record completeness at scale, grounding, and measured cost together.
Boyang Zhang, Adrian Lyjak, Elizabeth Stewart et al.· arXiv.org· 1 citation
This article proposes a formal representation of this type of document through a semantic description layer that formally describes the document's logical structure and captures its underlying semantics, and demonstrates the practical utility of these Ontological Cores through three distinct use cases.
M. El Ouaazizi· International Journal of Edu...· 0 citations
The problem of parsing heterogeneous PDF documents is still unresolved because of the use of complex layouts, embedded watermarks, and the disposition of the current systems to generate incoherent knowledge graphs (KGs) consisting of hundreds of nodes that are loosely connected and do not contribute a lot to education....
A production extraction layer that converts a live document stream into a validated knowledge graph aligned to a formal ontology, and improved search recall from roughly 70 to 95 percent with no false merges, and corrected seven classes of silent quality defect.
Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik· arXiv.org· 0 citations
A seven-stage graph-grounded pipeline that converts domain documents into a complete, auditable Web Ontology Language (OWL) Terminological Box (TBox) without any unconstrained generation step is presented, demonstrating that the pipeline produces stable, reusable domain representations from large document corpora.
Maruf Ahmed Mridul, A. Talukder, O. Seneviratne· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.