MQM-BR: A Framework for Assessing the Quality of Brazilian Open Data Metadata
Abstract
Open data initiatives depend on high-quality metadata to ensure that datasets can be discovered, understood, integrated, and reused across governmental, academic, and societal contexts. However, many public data portals still lack systematic and interoperable mechanisms to assess the structural and semantic quality of their metadata. This paper presents MQM-BR, a methodological framework for metadata quality assessment in the Brazilian open data ecosystem. The approach integrates FAIR-aligned metrics, semantic validation using SHACL, and a semantic representation of quality measurements based on an extension of the W3C Data Quality Vocabulary (DQV). The evaluation presented in this paper focuses on the quantitative component of MQM-BR, while the AI-assisted qualitative layer is presented as an ongoing component intended to translate diagnostic results into actionable recommendations. To evaluate the approach, we analyzed 5,000 datasets from the Brazilian Open Data Portal (dados.gov.br), converting catalog metadata into RDF according to the DCAT-BR profile and applying automated validation and scoring procedures. The results reveal significant gaps in metadata quality, particularly in contextual and interoperability-related attributes such as spatial and temporal coverage, controlled vocabularies, and semantic compliance. The proposed framework provides a reproducible and interoperable methodology for monitoring and diagnosing metadata quality, while establishing the foundation for future AI-assisted metadata improvement, contributing to stronger data governance practices and increased transparency, interoperability, and reuse in open government data ecosystems.