A dual-branch framework improves image quality by integrating unique features from infrared and visible data.
Infrared and Visible Image Fusion (IVIF) aims to integrate complementary multi-modal data to create a high-quality image that retains key information from both modalities. Current techniques face challenges in effectively distinguishing the features unique to each image type. Meanwhile, the attention mechanisms show limitations when integrating features from both the spatial and channel dimensions simultaneously. To tackle these challenges, we propose a dual-branch fusion framework called Cross-modal Differential and Dual-axis Attention Network (CDDANet). It consists of two branches, namely Global Enhance (GE) branch and Multi-scale Enhance (ME) branch. The GE branch leverages global information to enhance cross-modal features, while the ME branch is capable of processing information at multiple spatial scales. To boost the model's capacity to focus on key spatial and channel dimensions, a Dual-Axis Attention (DA) module is designed. Meanwhile, the Cross-modal Differential Attention (CDA) module is introduced to capture the differential features and the unique information between the two modalities. Comprehensive experiments on public datasets show that the proposed framework delivers strong performance. Both qualitative and quantitative evaluations show that it outperforms existing state-of-the-art methods. The code is available at https://github.com/lishuohui123/CDDANet.
No takes yet. Share an insight, caveat, or question.
Li et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: