Encoding, Linking, Retrieving: A Methodological Framework for Knowledge-Enriched Interview Corpora
-
University of Belgrade, Faculty of Mining and Geology,
Djusina 7, Belgrade, Serbia
ranka.stankovic@rgf.bg.ac.rs (corresponding author), biljana.rujevic@rgf.bg.ac.rs -
University Metropolitan, Faculty of Management
Tadeuša Košćuška 63, Belgrade, Serbia
tamara.vucenovic@metropolitan.ac.rs -
University of Belgrade, Faculty of Philology
Studentski Trg 2, Belgrade, Serbia
milica.ikonic.nesic@fil.bg.ac.rs
Abstract
This paper presents an AI-driven pipeline for transforming the interviews from “Digitalne Ikone 20+” book from unstructured transcripts into a structured, semantically enriched, and queryable knowledge resource. The raw text was first converted into XML-TEI format, with explicit structural markup of interview boundaries, speaker turns, paragraphs, temporal metadata, and topics. This encoding established logical segmentation and enabled targeted queries, such as retrieving content by speaker or thematic segment. An NLP and textometric analysis was conducted using the TXM tool and JeRTeh resources, followed by automatic Named Entity Recognition (NER) using models from the TESLA project. Key entity types were identified and embedded into the TEI structure. In the subsequent Named Entity Linking (NEL) stage, entities were disambiguated and connected to Wikidata identifiers, enriching the corpus with external knowledge graph references. Missing entities were added to Wikidata, contributing new structured knowledge. The resulting resource allows researchers, students, and the public to explore cultural heritage interviews through intelligent querying, automated dataset generation, and knowledge graph integration. The pipeline offers a replicable methodology for converting oral archives into AI-accessible knowledge bases.
Key words
Corpus, Named Entity Recognition, Named Entity Linking, Information Extraction, Textometry, Knowledge Graphs
Digital Object Identifier (DOI)
https://doi.org/10.2298/CSIS260523042S
Publication information
Volume 23, Issue 4 (September 2026)
Year of Publication: 2026
ISSN: 2406-1018 (Online)
Publisher: ComSIS Consortium
Full text
Available in PDF
Portable Document Format
How to cite
Stanković, R., Vučenović, T., Ikonić Nešić, M., Rujević, B.: Encoding, Linking, Retrieving: A Methodological Framework for Knowledge-Enriched Interview Corpora. Computer Science and Information Systems, 23(4) (2026). https://doi.org/10.2298/CSIS260523042S
Journal's Facebook page