Journal of Jilin University(Engineering and Technology Edition) ›› 2026, Vol. 56 ›› Issue (9): 2435-2443.doi: 10.13229/j.cnki.jdxbgxb.20250159

Previous Articles    

RGB⁃D semantic segmentation algorithm based on feature fusion attention

Sheng JIANG1(),Qi LU1,Miao-lei XIA2(),Jia-yu WANG1   

  1. 1.College of Physics,Changchun University of Science and Technology,Changchun 130022,China
    2.College of Architecture and Energy Engineering,Wenzhou University of Technology,Wenzhou 325035,China
  • Received:2025-02-28 Online:2026-09-01 Published:2026-09-07
  • Contact: Miao-lei XIA E-mail:js1985_cust@163.com;240654931@qq.com

Abstract:

A multimodal semantic segmentation model was proposed for indoor scene perception tasks. Existing multimodal fusion methods mainly focus on a single modality, such as RGB or depth image attention, while lacking effective utilization of attention in the joint feature space after modality fusion. Additionally, many networks suffer from a significant drop in accuracy when the number of parameters is reduced, due to their large parameter sizes. The RGB and depth features are extracted using ConvNeXt and MiT, respectively. A Cross-modal Feature Correction Module (CMFCM) is introduced to enhance spatial and channel features, facilitating interaction and correction between modalities. Additionally, a Cross-modal Attention Fusion Module (CMAFM) is designed to apply attention in the joint feature space, enabling effective multi-scale feature fusion and maximizing the utilization of segmentation information from both modalities. Depthwise separable convolution is also employed to significantly reduce computational costs. Finally, experiments on the NYU-Depth V2 dataset show that, compared with CMX, Dformer, and SA-Gate models, the proposed method improves segmentation accuracy by 1% while reducing the number of parameters by 15.1%, significantly lowering computational complexity.

Key words: semantic segmentation, multimodal fusion, indoor perception, attention mechanism

CLC Number: 

  • TP301.6

Fig.1

Diagram of model framework"

Fig.2

Depth feature extraction module"

Fig.3

Cross-modal Feature Correction Module"

Fig.4

Cross-modal attention fusion module"

Fig.5

Channel rearrangement"

Fig.6

Depth-separable convolution"

Table 1

Comparative experimental results"

方法PA/%MIoU/%参数量/M
CMX-B282.849.767
Dformer82.752.053
SA-Gate82.549.465
本文83.752.845

Fig.7

Comparison of visual segmentation results"

Table 2

Ablation results"

CMFCMCMAFM输入PA/%MIoU/%
××RGB72.147.6
××RGB-D74.748.5
×RGB-D76.649.8
×RGB-D77.250.5
RGB-D83.752.8
[1] 王文俊. 基于深度卷积神经网络的点云数据语义分割方法研究[D].成都:电子科技大学信息与通信工程学院, 2024.
Wang Wen-jun. Research on semantic segmentation of point cloud data based on deep convolutional neural networks[D].Chengdu: School of Information and Communication Engineering, University of Electronic Science and Technology of China, 2024.
[2] Jiang Jin-dong, Zheng Lu-nan, Luo Fei, et al. Rednet: residual encoder-decoder network for indoor RGB-D semantic segmentation[J/OL]. [2025-02-10]. .
[3] Seichter D, Köhler M, Lewandowski B, et al. Efficient RGB-D semantic segmentation for indoor scene analysis[C]∥2021 IEEE International Conference on Robotics and Automation, Xi'an, China, 2021: 13525-13531.
[4] Du S Q, Tang S J, Wang W X, et al. PSCNET: efficient RGB-D semantic segmentation parallel network based on spatial and channel attention[J]. ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, 2022, 1: 129-136.
[5] Chen X, Lin K Y, Wang J, et al. Bi-directional cross-modality feature propagation with separation-and-aggregation gate for RGB-D semantic segmentation[C]∥European Conference on Computer Vision, Glasgow, UK, 2020: 561-577.
[6] Zhang Jia-ming, Liu Hua-yao, Yang Kai-lun, et al. CMX: cross-modal fusion for RGB-X semantic segmentation with transformers[J]. IEEE Transactions on Intelligent Transportation Systems, 2023,24(12): 14679-14694.
[7] Yin Bo-wen, Zhang Xu-ying, Li Zhong-yu, et al. Dformer: rethinking RGBD representation learning for semantic segmentation[J/OL].[2025-02-10]. .
[8] Bui M, Alexis K. Diffusion-based RGB-D semantic segmentation with deformable attention transformer[J/OL].[2025-02-10]. .
[9] Dong S, Feng Y, Yang Q, et al. Efficient multimodal semantic segmentation via dual-prompt learning[C]∥2024 IEEE/RSJ International Conference on Intelligent Robots and Systems, Abu Dhabi, UAE, 2024: 14196-14203.
[10] 缪君, 严杰, 杜荣华, 等.基于双向特征融合的物体位姿估计方法[J].吉林大学学报: 工学版, 2026, 56(2): 523-532.
Miao Jun, Yan Jie, Du Rong-hua, et al. Object pose estimation method based on bidirectional feature fusion[J]. Journal of Jilin University (Engineering and Technology Edition), 2026, 56(2): 523-532.
[11] 王雪, 李占山, 吕颖达.基于多尺度感知和语义适配的医学图像分割算法[J].吉林大学学报: 工学版,2022, 52(3): 640-647.
Wang Xue, Li Zhan-shan, Ying-da Lyu. Medical image segmentation algorithm based on multi-scale perception and semantic adaptation[J]. Journal of Jilin University (Engineering and Technology Edition),2022, 52(3): 640-647.
[12] 周大可, 张超, 杨欣. 基于多尺度特征融合及双重注意力机制的自监督三维人脸重建[J].吉林大学学报: 工学版, 2022, 52(10): 2428-2437.
Zhou Da-ke, Zhang Chao, Yang Xin. Self-supervised 3D face reconstruction based on multi-scale feature fusion and dual attention mechanism[J]. Journal of Jilin University (Engineering and Technology Edition), 2022, 52(10): 2428-2437.
[13] Xie E, Wang W, Yu Z, et al. SegFormer: simple and efficient design for semantic segmentation with transformers[J]. Advances in Neural Information Processing Systems, 2021, 34: 12077-12090.
[14] Liu Z, Mao H, Wu C Y, et al. A convnet for the 2020s[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, USA, 2022: 11976-11986.
[15] Song Y, Wen J, Liu D, et al. Deep robotic grasping prediction with hierarchical RGB-D fusion[J]. International Journal of Control, Automation and Systems, 2022, 20(1): 243-254.
[16] Liu Z, Tan Y, He Q, et al. SwinNet: swin transformer drives edge-aware RGB-D and RGB-T salient object detection[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2021, 32(7): 4486-4497.
[17] 钱白云, 吕朝阳, 张维宁, 等. 基于多传感器信息融合与混合感受野残差卷积神经网络的调相机转子故障诊断[J].计算机测量与控制, 2023, 31(9): 29-35.
Qian Bai-yun, Lv Chao-yang, Zhang Wei-ning, et al. Phase condenser rotor fault diagnosis based on multi-sensor information fusion and mixed receptive field residual convolutional neural network[J]. Computer Measurement & Control, 2023, 31(9): 29-35.
[18] 孙启超, 恩擎, 段立娟, 等. 基于多模态自适应卷积的RGB-D图像语义分割[J]. 计算机辅助设计与图形学学报, 2022, 34(8): 1272-1282.
Sun Qi-chao, Qing En, Duan Li-juan, et al. RGB-D image semantic segmentation based on multimodal adaptive convolution[J]. Journal of Computer-Aided Design & Graphics, 2022, 34(8): 1272-1282.
[1] Xin-hui LIU,Zhuo-qun CHEN,Yan LYU. Review of industrial fault diagnosis based on deep learning [J]. Journal of Jilin University(Engineering and Technology Edition), 2026, 56(7): 1759-1779.
[2] Jie CAO,Zhi-feng CHEN,Jin-hua WANG,Li CHEN. Fault diagnosis method for gearbox with few samples based on diffusion model and DenseNet [J]. Journal of Jilin University(Engineering and Technology Edition), 2026, 56(7): 1787-1797.
[3] Feng SHI,Peng NIU,Min FAN. Uneven deformation detection of highway subgrade and pavement based on Faster R-CNN algorithm [J]. Journal of Jilin University(Engineering and Technology Edition), 2026, 56(7): 1950-1957.
[4] Jun MIAO,Jie YAN,Rong-hua DU,Lei LI,Jun CHU. A bidirectional feature fusion method for object position estimation [J]. Journal of Jilin University(Engineering and Technology Edition), 2026, 56(2): 523-532.
[5] Qiu-zhan ZHOU,Xin-meng LI,Hao-qing-zi SHEN,Hui-nan WU,Yuan-yuan LI,Jing RONG,Chun-hua HU,Ping-ping LIU. Non-intrusive load decomposition of unbalanced data based on attention mechanism [J]. Journal of Jilin University(Engineering and Technology Edition), 2026, 56(1): 239-246.
[6] Zhi-gang FENG,Meng-yuan REN,Bing DONG,Ming-yue YU. Rolling bearing fault diagnosis based on multi-band feature map and improved SqueezeNet [J]. Journal of Jilin University(Engineering and Technology Edition), 2026, 56(1): 96-108.
[7] Zhen HUO,Li-sheng JIN,Qiang HUA, HEYang. Edge feature⁃guided semantic segmentation method for intelligent vehicle [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(9): 3032-3041.
[8] Qing-lin AI,Yuan-xiao LIU,Jia-hao YANG. Small target swmantic segmentation method based MFF-STDC network in complex outdoor environments [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(8): 2681-2692.
[9] Yan PIAO,Ji-yuan KANG. RAUGAN:infrared image colorization method based on cycle generative adversarial networks [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(8): 2722-2731.
[10] Shan-na ZHUANG,Jun-shuai WANG,Jing BAI,Jing-jin DU,Zheng-you WANG. Video-based person re-identification based on three-dimensional convolution and self-attention mechanism [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(7): 2409-2417.
[11] Zhi-gang FENG,Shou-qi WANG,Ming-yue YU. Rolling bearing fault diagnosis based on variational mode extraction and lightweight network [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(6): 1883-1891.
[12] Ying YU,Chun-ping WANG,Ren-ke KOU,Bo-xiong YANG,Lei WANG,Fu-jun ZHAO,Qiang FU. Semantic segmentation algorithm for multi temporal high⁃resolution satellite remote sensing images [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(6): 2131-2137.
[13] Ya-li XUE,Tong-an YU,Shan CUI,Li-zun ZHOU. Infrared small target detection based on cascaded nested U-Net [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(5): 1714-1721.
[14] He-shan ZHANG,Meng-wei FAN,Xin TAN,Zhan-ji ZHENG,Li-ming KOU,Jin XU. Dense small object vehicle detection in UAV aerial images using improved YOLOX [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(4): 1307-1318.
[15] Hua CAI,Yu-yao WANG,Qiang FU,Zhi-yong MA,Wei-gang WANG,Chen-jie ZHANG. Semantic segmentation network based on attention mechanism and feature fusion [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(4): 1384-1395.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!