LLMCrater is a lifecycle-aware metadata generation framework that combines Large Language Models (LLMs) with stage-specific RO-Crate metadata profiles that progressively enriches metadata across four research lifecycle stages while remaining compatible with RO-Crate~1.1 and EOSC metadata recommendations.
Abstract
FAIR (Findable, Accessible, Interoperable, and Reusable) metadata is essential for the discovery, interoperability, and reuse of scientific research assets. However, creating and maintaining FAIR metadata remains largely manual, making the process time-consuming for heterogeneous research artifacts generated throughout the research lifecycle. Existing approaches primarily generate metadata at publication time, missing opportunities to capture contextual information as it becomes available. To address this limitation, we present \emph{LLMCrater}, a lifecycle-aware metadata generation framework that combines Large Language Models (LLMs) with stage-specific RO-Crate metadata profiles. The framework progressively enriches metadata across four research lifecycle stages (Design, Development, Deployment, and Execution \&Provenance) while remaining compatible with RO-Crate~1.1 and EOSC metadata recommendations. It automatically extracts metadata from heterogeneous artifacts, generates and validates machine-actionable RO-Crates, and supports publication to FAIR repositories and PID services (e.g., Zenodo). We demonstrate the approach using two representative use cases: a 5G experimentation environment within SLICES-RI and an experiment on GreenDIGIT's EcoJupyter platform. Results show that LLMCrater progressively enriches metadata throughout the research lifecycle and generates valid RO-Crates conforming to the RO-Crate~1.1 specification.
This study contributes a systematic workflow for metadata mapping and provides empirical evidence on the limitations of current metadata standardization practices, supporting future efforts toward improved cross-domain metadata interoperability.
Yan Cong, Masao Takaku, Yasuyuki Minamiyama et al.· 0 citations
It is worthwhile to make institutional repositories interoperable with Web-based computational reproducibility tools to foster the reuse of scholarship. While repositories excel at long-term preservation, they were not made with the interactive capabilities required to execute complex code or analyze dynamic data direc...
The design of a workflow blueprint compliant with FAIR principles and grounded in Open Science practices, developed for the retrieval, harmonisation, and publication of curated bibliographic metadata is introduced.
Arianna Moretti, Iiro Tiihonen, Jonas Fischer· 0 citations
This work investigates the feasibility of using Generative AI to assist human operators working for data producers in creating structured dataset descriptions compliant with GeoDCAT-AP standards, and proposes a six-layer prompting framework and evaluates seven ablation strategies.
M. Di Leo, Ilyas Tiouassiouine-Maes, Jordi Escriu et al.· Open Research Europe· 0 citations
LLMVul, a vulnerability-labeled dataset of LLM-generated C/C++ functions mined from real production repositories, enables reproducible research on vulnerability detection, security evaluation of LLM-generated code, and analysis of vulnerability patterns in AI-assisted software development.
The results show that structured RDF pipelines currently provide the most stable integration behavior, whereas JSON and text pipelines remain more sensitive to errors in mapping, extraction, and linking.
Marvin Hofer, Erhard Rahm· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.