Decoupled Global-Local Collaborative Network for Visual Question Answering

Ganlong Zhou1, Dezhi Han2, Xiang Shen1,3, Xiaoqi Xin1

  1. School of Information Engineering, Shanghai Maritime University
    Shanghai 201306, China
    ganlongzhou@163.com (corresponding author)
  2. School of Information Engineering, Shanghai Maritime University
    Shanghai 201306, China
    dzhan@shmtu.edu.cn (corresponding author)
  3. School of Computer Science, The University of Sydney
    NSW 2006, Sydney, Australia
    xiang.shen@sydney.edu.au

Abstract

Visual Question Answering (VQA) aims to achieve cross-modal semantic understanding through joint modeling of visual content and natural language. Although existing attention-based approaches effectively align features, they struggle to simultaneously accommodate global semantic modeling and local fine-grained perception within single-path architectures, leading to semantic drift and underutilization of information when addressing multi-perspective reasoning tasks. To mitigate this limitation, we propose a Decoupled Global-Local Collaborative Network (DGL-Net) that structurally partitions the cross-modal reasoning process into two complementary branches: global semantic understanding and local fine-grained modeling. Specifically, the global branch employs hierarchical attention fusion to model long-range dependencies and holistic scene semantics, providing contextual constraints for reasoning. The local branch introduces semantic relation enhancement and relevant region screening to explicitly model intra-modal structural relationships, strengthening fine-grained region-word alignment by injecting structural information. During training and inference, a branch-decoupling and collaborative fusion strategy enables independent yet complementary feature learning. Extensive experimental results demonstrate accuracies of 72.15% and 57.87% on VQA-v2 and GQA benchmarks, respectively. Notably, the proposed method exhibits enhanced robustness and generalization in complex semantic and compositional reasoning tasks, validating the efficacy of the decoupled collaborative strategy.

Key words

Visual Question Answering, Transformer, Decoupled Strategy, Cross-Modal Alignment

Digital Object Identifier (DOI)

https://doi.org/10.2298/CSIS260403039Z

Publication information

Volume 23, Issue 4 (September 2026)
Year of Publication: 2026
ISSN: 2406-1018 (Online)
Publisher: ComSIS Consortium

Full text

Download Available in PDF
Portable Document Format

How to cite

Zhou, G., Han, D., Shen, X., Xin, X.: Decoupled Global-Local Collaborative Network for Visual Question Answering. Computer Science and Information Systems, 23(4) (2026). https://doi.org/10.2298/CSIS260403039Z