Research Progress of Multimodal Emotion Recognition Fusion Strategy
DOI:
https://doi.org/10.54097/rw0rnn76Keywords:
Multimodal Sentiment Recognition; Fusion Strategies; Cross-modal Attention; LLM-Centric Fusion; Lightweight Models.Abstract
Multimodal emotion recognition relies on the collaborative perception of text, speech, visual and physiological signals. In recent years, it has achieved rapid development in the fields of emotional computing, human-machine interaction and medical health. With the emergence of large models, cross-modal attention mechanisms and dynamic fusion strategies, traditional decision-level fusion and feature-level fusion are unable to meet the requirements of refined, real-time and open vocabulary recognition. This article reviews five cutting-edge fusion methods, namely model-level fusion, cross-modal attention, multimodal fusion centered on large language models, dynamic fusion driven by reinforcement learning, and lightweight models for deployment. The article discusses the core concepts, advantages and disadvantages, and applicable scenarios of each method. Finally, it identifies current challenges and outlines future research directions.
Downloads
References
[1] Gajiwala, C. The Rise of Deep Learning and Neural Networks: Revolutionizing Artificial Intelligence. European Journal of Computer Science and Information Technology,2025.
[2] Chen, Z., & Zhang, B. Application of Multimodal Deep Learning Algorithm in Image and Text Fusion Recognition. 2025 IEEE 5th International Conference on Power, Electronics and Computer Applications, 2025, 599-604.
[3] C. Xu, H. Zhao, X. Lu, et al. A Deep Reinforcement Learning Method for Autonomous Driving Integrating Multi-Modal Fusion, IEEE Transactions on Intelligent Transportation Systems, 2025, 26(8): 11850-11863.
[4] Geetha A.V., Mala T., Priyanka D., et al. Multimodal Emotion Recognition with Deep Learning: Advancements, challenges, and future directions. Information Fusion, 2024, 105.
[5] Magauiya Zhussip, Dmitriy Shopkhoev, Ammar Ali, et al. Share Your Attention: Transformer Weight Sharing via Matrix - based Dictionary Learning. arXiv, 2025, 2508.04581
[6] R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, IEEE Transactions on Neural Networks, 1998, 9(5): 1054-1054
[7] Puterman, M. Markov Decision Processes: Discrete Stochastic Dynamic Programming. IMA Journal of Management Mathematics, 1994.
[8] T. Xu, Y. Pang, Y. Zhu, et al. Real-Time Driving Style Integration in Deep Reinforcement Learning for Traffic Signal Control, IEEE Transactions on Intelligent Transportation Systems,2025, 26(8): 11879-11892.
[9] Z. Zhang, F. Liang, W. Wang, et al. Skeleton-Based Pretraining with Discrete Labels for Emotion Recognition in IoT Environments, IEEE Internet of Things Journal, 2025, 12(15): 31856-31868.
[10] Ashutosh Holla B., Manohara Pai M.M., Ujjwal Verma, et al. MSFFT: Multi - Scale Feature Fusion Transformer for cross platform vehicle re - identification. Neurocomputing, 2024, 582.
[11] Lai, D., Zhang, Y., Liu, Y., et al. Deep Learning - Based Multi - Modal Fusion for Robust Robot Perception and Navigation. arXiv, 2025, 2504.19002.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Highlights in Science, Engineering and Technology

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.







