This article examines how TEI-based digital archives and human-curated annotation layers can support documented uses of large language models in literary analytics and proposes a layered workflow in which TEI, taxonomies, authority records, annotation tables, machine-readable metadata, and validation procedures interact.
Abstract
This article examines how TEI-based digital archives and human-curated annotation layers can support documented uses of large language models in literary analytics. It argues that structured textual infrastructures strengthen the conditions for evaluating and interpreting LLM outputs by making scholarly decisions explicit, inspectable, auditable, and contestable. Here, semantic data governance refers to provenance, documentation, controlled vocabularies, interpretive criteria, and critical evaluation. The case study is the LdoD Archive and its Virtual Edition “Philosophical Intertextuality,” built on TEI-encoded sources for Fernando Pessoa’s Livro do Desassossego, a fragmentary and posthumously assembled work. The virtual edition adds a human-authored semantic layer through taxonomies and tagging tools that connect fragments with philosophical categories and named thinkers. Building on this architecture, the article proposes a layered workflow in which TEI, taxonomies, authority records, annotation tables, machine-readable metadata, and validation procedures interact. A curated dataset derived from the virtual edition can support qualitative comparison between human annotations and model outputs, exploratory evaluation of omissions, unsupported associations, and interpretive divergences, and prompt design based on concise scholarly examples. The annotations function as situated scholarly claims, allowing interpretive plurality to coexist with stronger provenance, inspection, and critical assessment in transparent and reproducible conditions for future LLM-assisted literary analysis across different analytical contexts.
This paper proposes a human-centered workflow in which AI assists editorial practices through schema-constrained suggestions, consistency checks, and semantic pattern detection while preserving human authority over interpretative decisions.
PubLink is presented, a modular toolset designed to bridge gaps in digital publishing by connecting existing systems and standards rather than replacing them, minimizing dependency on any single platform.
E. Bastianello, C. Tomlinson· Proceedings of the 37th ACM...· 0 citations
ReSB2 is introduced, a framework for retrieving and linking similar legislative bills that supports human–machine collaboration and helps reduce redundancy in the lawmaking process, and fine-tunes two ModernBERT-based language models on authentic legislative data.
Lucas G. L. Costa, Átila Souza, Elves Rodrigues et al.· Proceedings of the 37th ACM...· 0 citations
It is suggested that LLMs can serve as auxiliary annotators for sensitive-language detection in historical materials, provided that human oversight and contextual interpretation remain central to annotation workflows.
Ya-Hui Zhao, Clemencia Siro, L. Hollink et al.· 0 citations
This work targets a KG for Sophocles’ Antigone that supports two coupled uses: structured retrieval, through integrity and competency questions expressed in SPARQL over dramatic structure and interpretive annotations; and interactive exploration, through a lightweight read client that navigates lines across languages,...
Local governments are beginning to use artificial intelligence (AI) in internal administrative work, including recruitment. In a decentralized municipality, however, a standardized AI assessment must travel between a technology vendor, central HR specialists, recruiters, and hiring managers responsible for different se...
Jonny Holmström· Government Information Quart...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.