Skip to content
Open access

TEI, Taxonomies, and Semantic Data Governance for Documented LLM-Assisted Literary Analytics

Sep 2026 · Analytics · 0 citations · 17 references

TL;DR

This article examines how TEI-based digital archives and human-curated annotation layers can support documented uses of large language models in literary analytics and proposes a layered workflow in which TEI, taxonomies, authority records, annotation tables, machine-readable metadata, and validation procedures interact.

Abstract

This article examines how TEI-based digital archives and human-curated annotation layers can support documented uses of large language models in literary analytics. It argues that structured textual infrastructures strengthen the conditions for evaluating and interpreting LLM outputs by making scholarly decisions explicit, inspectable, auditable, and contestable. Here, semantic data governance refers to provenance, documentation, controlled vocabularies, interpretive criteria, and critical evaluation. The case study is the LdoD Archive and its Virtual Edition “Philosophical Intertextuality,” built on TEI-encoded sources for Fernando Pessoa’s Livro do Desassossego, a fragmentary and posthumously assembled work. The virtual edition adds a human-authored semantic layer through taxonomies and tagging tools that connect fragments with philosophical categories and named thinkers. Building on this architecture, the article proposes a layered workflow in which TEI, taxonomies, authority records, annotation tables, machine-readable metadata, and validation procedures interact. A curated dataset derived from the virtual edition can support qualitative comparison between human annotations and model outputs, exploratory evaluation of omissions, unsupported associations, and interpretive divergences, and prompt design based on concise scholarly examples. The annotations function as situated scholarly claims, allowing interpretive plurality to coexist with stronger provenance, inspection, and critical assessment in transparent and reproducible conditions for future LLM-assisted literary analysis across different analytical contexts.

Read PDF

Similar papers

Book Open access Sep 2026

ReSB²: Retrieving Similar Brazilian State Bills

ReSB2 is introduced, a framework for retrieving and linking similar legislative bills that supports human–machine collaboration and helps reduce redundancy in the lawmaking process, and fine-tunes two ModernBERT-based language models on authentic legislative data.

Lucas G. L. Costa, Átila Souza, Elves Rodrigues et al. · 0 citations
Review

Sophocles’ Antigone as a Knowledge Graph through a Hybrid Collaborative Workflow with Ontology-Guided LLM Extraction

This work targets a KG for Sophocles’ Antigone that supports two coupled uses: structured retrieval, through integrity and competency questions expressed in SPARQL over dramatic structure and interpretive annotations; and interactive exploration, through a lightweight read client that navigates lines across languages,...

Apostolos Baniotis, Marsel Senka, Entisa Tzeortziana Komoritsan et al. · 0 citations
Open access Sep 2026

Working with semantic machines: The role of semantic reconciliation in public sector data governance

Local governments are beginning to use artificial intelligence (AI) in internal administrative work, including recruitment. In a decentralized municipality, however, a standardized AI assessment must travel between a technology vendor, central HR specialists, recruiters, and hiring managers responsible for different se...

Jonny Holmström · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.