Semantic Labour and the Politics of Metadata
Abstract
The structured, machine-readable metadata that researchers and librarians generate to make scholarly outputs retrievable has become, in the age of artificial intelligence (AI), the decisive input shaping what knowledge can be found, ranked, and generated. This article argues that this metadata is produced through a cycle of semantic labourーthe cognitive, administrative, and technical work of encoding research outputs in structured, machine-readable formーthat is largely unremunerated by its commercial beneficiaries and that this labour is the mechanism through which a broader process of epistemic capture operates. As enriched metadata is enclosed within proprietary knowledge graphs and AI products, especially the retrieval-augmented generation (RAG) systems now central to scholarly search, the coverage gaps, classification biases, and ranking logics embedded in those systems are operationalised at scale, transferring authority over what counts as legible knowledge from scholarly communities to commercial infrastructural providers. This article traces this dynamic through a six-stage model of the metadata production cycle, analyses the mechanisms of commercial enclosure and their epistemic consequences, and assesses the conditions under which open metadata infrastructures could constitute a genuine counterweight. The conclusion argues that metadata sovereignty, understood as the capacity of scholarly communities to maintain collective authority over the semantic conditions of their epistemic legibility, is a central challenge for scholarly communication in the AI era.