SAMS-M: Explicit Sentiment-Guided Alignment and Multi-dimensional Mutual Supervision for Multimodal Sentiment Analysis
-
School of Automotive and Traffic Engineering, Tianjin University of Technology and Education
300222 Tianjin, China
{0721241391,0722231050}@tute.edu.cn -
School of Computer Science and Engineering, Anhui University of Science and Technology
232000 Huainan, China
2023201230aust.edu.cn, glzhu@aust.edu.cn (corresponding author) -
School of Mathematics and Big Data, Anhui University of Science and Technology
232000 Huainan, China
2025201731@aust.edu.cn, rob5li@outlook.com
Abstract
Multimodal Sentiment Analysis leverages the fusion of heterogeneous data to achieve fine-grained emotional understanding, which finds extensive application in large-scale public opinion monitoring and data mining. However, existing methods face two key challenges: (1) cross-modal alignment suffers from redundancy and semantic drift without explicit modeling of sentiment-critical cues, inducing spurious correlations; and (2) heterogeneous representation spaces lead to imbalanced modality contributions, particularly under weak image–text correlation or sentiment inconsistency. To address these challenges, we propose an explicit sentiment-guided alignment and multi-dimensional cross-modal mutual supervision-based model for multimodal sentiment analysis. The model primarily employs a fine-grained sentiment–saliency directed alignment mechanism, which leverages bidirectional cross-attention to couple textual sentiment cues with visual saliency, enabling precise localization of sentiment-relevant regions. Furthermore, we introduce a tripartite strong contrastive learning strategy to mitigate distribution discrepancies between heterogeneous modalities within a shared latent space, thereby enhancing cross-modal coherence and complementarity. Finally, we design a noise-robust gating-based fusion module, which, together with text augmentation and deep supervision, facilitates effective joint optimization. Experimental results show that SAMS-M obtains the best results on MVSA-Single and MSD and remains competitive on the noisier MVSA-Multiple benchmark; thus, the evidence supports strong but dataset-dependent performance rather than uniform state-of-the-art superiority.
Key words
Multimodal Sentiment Analysis, Cross-modal Alignment, Contrastive Learning, Sentiment-guided, Noise-robust Fusion
Digital Object Identifier (DOI)
https://doi.org/10.2298/CSIS260804043Q
Publication information
Volume 23, Issue 4 (September 2026)
Year of Publication: 2026
ISSN: 2406-1018 (Online)
Publisher: ComSIS Consortium
Full text
Available in PDF
Portable Document Format
How to cite
Qi, S., Zhang, Y., Zhang, S., Xu, J., Li, K.C., Zhu, G.: SAMS-M: Explicit Sentiment-Guided Alignment and Multi-dimensional Mutual Supervision for Multimodal Sentiment Analysis. Computer Science and Information Systems, 23(4) (2026). https://doi.org/10.2298/CSIS260804043Q
Journal's Facebook page