A framework for the automated thematic annotation of Open Government Data

Abdul Aziz1, Mohsan Ali2, Dagoberto José Herrera-Murillo1   , Maria Ioanna Maratsi2, Francisco J. Lopez-Pellicer1, Javier Nogueras-Iso1   

  1. Aragon Institute of Engineering Research (I3A), Universidad de Zaragoza, Zaragoza, Spain
    {abdul.aziz, dherrera, fjlopez}@unizar.es
    jnog@unizar.es (corresponding author)
  2. Department of Information and Communication Systems Engineering, University of the Aegean, Samos, Greece
    {mohsan, ioanna.m}@aegean.gr

Abstract

Governmental policies for transparency and reuse of public sector information have encouraged the launch of open government data portals around the world. Many of these portals are based on pyramidal structures: national open data portals are aggregators of the contents harvested from open data portals maintained by governments in charge of administrative areas with a narrower scope. Taking into account this hierarchical organization, these open data portals lack consistent and scalable mechanisms for thematic annotation, limiting dataset discoverability. This work proposes a framework for the automated thematic classification of open government data. The framework integrates (i) thematic annotation quality assessment, (ii) supervised machine learning models trained on annotated metadata corpora, and (iii) embedding-based semantic similarity methods for theme assignment in the absence of reliable annotations. The framework is evaluated using 29,793 datasets from data.europa.eu, the European open data portal. Experimental results show that supervised models achieve high classification performance, with Support Vector Machines reaching an accuracy of 93.65%, while unsupervised embedding-based approaches achieve substantial semantic agreement with portal-assigned themes (74.56%) using transformer-based representations. These results demonstrate that the proposed framework enables scalable, consistent, and interoperable thematic annotation, offering both theoretical contributions to automated metadata enrichment and practical value for integration into large-scale open data portal infrastructures.

Key words

Data Annotation, Open Government Data, Open Data Portals, Automated Annotation, Thematic Annotation

Digital Object Identifier (DOI)

https://doi.org/10.2298/CSIS251029022A

Publication information

Volume 23, Issue 2 (April 2026)
Year of Publication: 2026
ISSN: 2406-1018 (Online)
Publisher: ComSIS Consortium

Full text

Download Available in PDF
Portable Document Format

How to cite

Aziz, A., Ali, M., Herrera-Murillo, D.J., Maratsi, M.I., Lopez-Pellicer, F.J., Nogueras-Iso, J.: A framework for the automated thematic annotation of Open Government Data. Computer Science and Information Systems, 23(2), 917–946 (2026). https://doi.org/10.2298/CSIS251029022A