Encoding, linking, retrieving: A methodological framework for knowledge-enriched interview corpora
This paper presents an AI-driven pipeline for transforming the interviews from “Digitalne Ikone 20+” book from unstructured transcripts into a structured, semantically enriched, and queryable knowledge resource. The raw text was first converted into XML-TEI format, with explicit structural markup of interview boundaries, speaker turns, paragraphs, temporal metadata, and topics. This encoding established logical segmentation and enabled targeted queries, such as retrieving content by speaker or thematic segment. An NLP and textometric analysis was conducted using the TXM tool and JeRTeh resources, followed by automatic Named Entity Recognition (NER) using models from the TESLA project. Key entity types were identified and embedded into the TEI structure. In the subsequent Named Entity Linking (NEL) stage, entities were disambiguated and connected to Wikidata identifiers, enriching the corpus with external knowledge graph references. Missing entities were added to Wikidata, contributing new structured knowledge. The resulting resource allows researchers, students, and the public to explore cultural heritage interviews through intelligent querying, automated dataset generation, and knowledge graph integration. The pipeline offers a replicable methodology for converting oral archives into AI-accessible knowledge bases.