SAMS-M: Explicit Sentiment-Guided Alignment and Multi-dimensional Mutual Supervision for Multimodal Sentiment Analysis

Shishu Qi1, Yulei Zhang2, Siyang Zhang3, Jia Xu1, Kuan-Ching Li3, Guangli Zhu2

  1. School of Automotive and Traffic Engineering, Tianjin University of Technology and Education
    300222 Tianjin, China
    {0721241391,0722231050}@tute.edu.cn
  2. School of Computer Science and Engineering, Anhui University of Science and Technology
    232000 Huainan, China
    2023201230aust.edu.cn, glzhu@aust.edu.cn (corresponding author)
  3. School of Mathematics and Big Data, Anhui University of Science and Technology
    232000 Huainan, China
    2025201731@aust.edu.cn, rob5li@outlook.com

Abstract

Multimodal Sentiment Analysis leverages the fusion of heterogeneous data to achieve fine-grained emotional understanding, which finds extensive application in large-scale public opinion monitoring and data mining. However, existing methods face two key challenges: (1) cross-modal alignment suffers from redundancy and semantic drift without explicit modeling of sentiment-critical cues, inducing spurious correlations; and (2) heterogeneous representation spaces lead to imbalanced modality contributions, particularly under weak image–text correlation or sentiment inconsistency. To address these challenges, we propose an explicit sentiment-guided alignment and multi-dimensional cross-modal mutual supervision-based model for multimodal sentiment analysis. The model primarily employs a fine-grained sentiment–saliency directed alignment mechanism, which leverages bidirectional cross-attention to couple textual sentiment cues with visual saliency, enabling precise localization of sentiment-relevant regions. Furthermore, we introduce a tripartite strong contrastive learning strategy to mitigate distribution discrepancies between heterogeneous modalities within a shared latent space, thereby enhancing cross-modal coherence and complementarity. Finally, we design a noise-robust gating-based fusion module, which, together with text augmentation and deep supervision, facilitates effective joint optimization. Experimental results show that SAMS-M obtains the best results on MVSA-Single and MSD and remains competitive on the noisier MVSA-Multiple benchmark; thus, the evidence supports strong but dataset-dependent performance rather than uniform state-of-the-art superiority.

Key words

Multimodal Sentiment Analysis, Cross-modal Alignment, Contrastive Learning, Sentiment-guided, Noise-robust Fusion

Digital Object Identifier (DOI)

https://doi.org/10.2298/CSIS260804043Q

Publication information

Volume 23, Issue 4 (September 2026)
Year of Publication: 2026
ISSN: 2406-1018 (Online)
Publisher: ComSIS Consortium

Full text

Download Available in PDF
Portable Document Format

How to cite

Qi, S., Zhang, Y., Zhang, S., Xu, J., Li, K.C., Zhu, G.: SAMS-M: Explicit Sentiment-Guided Alignment and Multi-dimensional Mutual Supervision for Multimodal Sentiment Analysis. Computer Science and Information Systems, 23(4) (2026). https://doi.org/10.2298/CSIS260804043Q