Encoding, Linking, Retrieving: A Methodological Framework for Knowledge-Enriched Interview Corpora

Ranka Stanković1, Tamara Vučenović2, Milica Ikonić Nešić3, Biljana Rujević1

  1. University of Belgrade, Faculty of Mining and Geology,
    Djusina 7, Belgrade, Serbia
    ranka.stankovic@rgf.bg.ac.rs (corresponding author), biljana.rujevic@rgf.bg.ac.rs
  2. University Metropolitan, Faculty of Management
    Tadeuša Košćuška 63, Belgrade, Serbia
    tamara.vucenovic@metropolitan.ac.rs
  3. University of Belgrade, Faculty of Philology
    Studentski Trg 2, Belgrade, Serbia
    milica.ikonic.nesic@fil.bg.ac.rs

Abstract

This paper presents an AI-driven pipeline for transforming the interviews from “Digitalne Ikone 20+” book from unstructured transcripts into a structured, semantically enriched, and queryable knowledge resource. The raw text was first converted into XML-TEI format, with explicit structural markup of interview boundaries, speaker turns, paragraphs, temporal metadata, and topics. This encoding established logical segmentation and enabled targeted queries, such as retrieving content by speaker or thematic segment. An NLP and textometric analysis was conducted using the TXM tool and JeRTeh resources, followed by automatic Named Entity Recognition (NER) using models from the TESLA project. Key entity types were identified and embedded into the TEI structure. In the subsequent Named Entity Linking (NEL) stage, entities were disambiguated and connected to Wikidata identifiers, enriching the corpus with external knowledge graph references. Missing entities were added to Wikidata, contributing new structured knowledge. The resulting resource allows researchers, students, and the public to explore cultural heritage interviews through intelligent querying, automated dataset generation, and knowledge graph integration. The pipeline offers a replicable methodology for converting oral archives into AI-accessible knowledge bases.

Key words

Corpus, Named Entity Recognition, Named Entity Linking, Information Extraction, Textometry, Knowledge Graphs

Digital Object Identifier (DOI)

https://doi.org/10.2298/CSIS260523042S

Publication information

Volume 23, Issue 4 (September 2026)
Year of Publication: 2026
ISSN: 2406-1018 (Online)
Publisher: ComSIS Consortium

Full text

Download Available in PDF
Portable Document Format

How to cite

Stanković, R., Vučenović, T., Ikonić Nešić, M., Rujević, B.: Encoding, Linking, Retrieving: A Methodological Framework for Knowledge-Enriched Interview Corpora. Computer Science and Information Systems, 23(4) (2026). https://doi.org/10.2298/CSIS260523042S