Skip to content
Open access

Towards AI-assisted metadata generation for improved description of geospatial data

Aug 2026 · Open Research Europe · Vol 6, pp. 277 · 0 citations · 35 references

TL;DR

This work investigates the feasibility of using Generative AI to assist human operators working for data producers in creating structured dataset descriptions compliant with GeoDCAT-AP standards, and proposes a six-layer prompting framework and evaluates seven ablation strategies.

Abstract

Metadata are fundamental components of spatial data infrastructures, enabling the discovery, evaluation, and reuse of datasets. Their creation and maintenance, typically performed manually, can be costly, time-consuming, and prone to inconsistencies. This work investigates the feasibility of using Generative AI (GenAI) to assist human operators working for data producers in creating structured dataset descriptions compliant with GeoDCAT-AP standards. We address the following research question: to what extent can layered prompt engineering strategies improve the quality of metadata descriptions generated through Large Language Models (LLMs), and how do individual prompt components-role definition, content rules, template structure, and few-shot examples-interact with model selection to affect structural compliance, semantic similarity, and factual reliability? We propose a six-layer prompting framework and evaluate seven ablation strategies, each selectively disabling specific layers, using two LLMs (qwen3-32b and qwen3-coder-30b-a3b-instruct) and eight geospatial datasets. Generated descriptions are assessed using eight automated metrics spanning lexical quality, structural compliance, and semantic similarity, complemented by targeted expert evaluation of factual accuracy. Results reveal a clear hierarchy of layer impact. The input template, used as the reference structure of the description, is the most influential component, driving both structural compliance and readability. Few-shot examples are the second most impactful layer, substantially reducing redundancy and improving semantic alignment. Content rules do not measurably improve surface-level output quality, but serve as critical safeguards, encouraging models to flag missing information rather than fabricate content. Expert role definition contributes the least measurable effect. Both LLMs produce semantically comparable outputs under full guidance but diverge without structural constraints. We recommend the full prompting configuration for production use and provide practical guidelines for balancing prompt complexity, output quality, and factual reliability in LLM-assisted metadata generation workflows.

Read PDF

Similar papers

Review Open access Sep 2026

Accelerating metadata annotation in collaborative research centers: A hybrid AI workflow for biomedical entities

Collaborative Research Centers rely on FAIR-compliant, richly structured metadata, yet manual annotation is a major bottleneck. We implemented a search-augmented large language model (LLM) workflow within a local research data management system to pre-annotate biomedical entities, using human-in-the-loop verifica...

M. Watter, F. Engel, Aref Kalantari et al. · 0 citations
Open access Aug 2026

Towards a Spatial Knowledge Mesh: A Metamodel for Federated and Interoperable Spatial Knowledge Graphs to Enable Geospatial Awareness, Integrity, Provenance, and Trust in Large Language Models

Abstract. Large language models (LLMs) have the potential to make geospatial decision support more accessible to non-technical users and under-resourced settings in public health, disaster response, and infrastructure planning. Compared with traditional geospatial workflows requiring specialized expertise, manual data...

N. McEachen, Justin Smethie · 0 citations
Preprint Aug 2026

LLMCrater: Lifecycle-Aware FAIR Metadata Generation using Large Language Models

LLMCrater is a lifecycle-aware metadata generation framework that combines Large Language Models (LLMs) with stage-specific RO-Crate metadata profiles that progressively enriches metadata across four research lifecycle stages while remaining compatible with RO-Crate~1.1 and EOSC metadata recommendations.

Dani Termaat, N. Soveizi, Zhi-Ming Zhao et al. · 0 citations

AI-Guided Metadata Construction for Meaning-Driven Digital Knowledge Systems: A Framework for Automated Metadata Generation and Semantic Discovery

An AI-guided framework is developed that aligns AI-assisted metadata extraction with Dublin Core Terms and the FAIR principles for digital libraries, archives, and cultural-heritage repositories and is evaluated as a design-science artefact in which retrieval is not a side feature but a feedback loop.

Wirapong Chansanam, Umawadee Detthamrong, Chunqiu Li et al. · 0 citations
Review Open access Aug 2026

Human-in-the-Loop Generative AI for KOS Registry Publishing: Designing a Dual-Automation Workflow

A practical, quality-assured method for publishing KOS-related RFPs as context-bearing registry records is contributed by integrating a two-track pipeline—scripted page generation for structured fields and generative summarization for narrative RFP text—with HITL governance for expert verification and correction.

Ziyoung Park · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.