Skip to content

Author

R. Stanković

We have 3 of 108 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access 2026

Encoding, linking, retrieving: A methodological framework for knowledge-enriched interview corpora

This paper presents an AI-driven pipeline for transforming the interviews from “Digitalne Ikone 20+” book from unstructured transcripts into a structured, semantically enriched, and queryable knowledge resource. The raw text was first converted into XML-TEI format, with explicit structural markup of interview boundaries, speaker turns, paragraphs, temporal metadata, and topics. This encoding established logical segmentation and enabled targeted queries, such as retrieving content by speaker or thematic segment. An NLP and textometric analysis was conducted using the TXM tool and JeRTeh resources, followed by automatic Named Entity Recognition (NER) using models from the TESLA project. Key entity types were identified and embedded into the TEI structure. In the subsequent Named Entity Linking (NEL) stage, entities were disambiguated and connected to Wikidata identifiers, enriching the corpus with external knowledge graph references. Missing entities were added to Wikidata, contributing new structured knowledge. The resulting resource allows researchers, students, and the public to explore cultural heritage interviews through intelligent querying, automated dataset generation, and knowledge graph integration. The pipeline offers a replicable methodology for converting oral archives into AI-accessible knowledge bases.

R. Stanković, Tamara Vučenović, Milica Ikonic-Nesic et al. · 0 citations
Open access 2026

PARSEME 2.0 Multilingual Corpus of Multiword Expressions

We present edition 2.0 of the PARSEME multilingual corpus annotated for multiword expressions (MWEs), resulting from efforts of the PARSEME community towards universality-driven modeling of idiomaticity. With respect to previous editions, we extend the annotation scope to all syntactic MWE categories: verbal, nominal, adjectival, adverbial and functional. We cover 17 languages, of which 7 are new. The annotation process is based on cross-lingually unified guidelines, phrased as decision diagrams over linguistic tests, and a typology of 18 MWE categories. The corpus contains almost 5 million tokens, over 250,000 sentences and 140,000 MWE annotations. The applicability of the corpus is tested in baseline experiments with a prompt-based MWE identification system. Results show that generic large language models do not encode sufficient knowledge to solve the MWE identification task.

Agata Savary, Manon Scholivet, Carlos Ramisch et al. · 1 citation
Aug 2026

A semi-automated LLM-based framework for word sense disambiguation in Serbian

LLM-assisted sense assignment with a Serbian WordNet-based custom inventory, iterative inventory expansion, and expert validation is combined with a constrained JSON-formatted output to support the practical construction and refinement of sense-annotated resources in a low-resource setting.

Saša Petalinkar, R. Stanković, Milica Ikonić Nešić et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.