吉林大学学报(理学版) ›› 2026, Vol. 64 ›› Issue (4): 823-0834.

• • 上一篇    下一篇

基于特征增强的3D目标检测模型ICasA

贺怀清, 翟羽佳, 刘浩翰, 惠康华   

  1. 中国民航大学 计算机与人工智能学院, 天津 300300
  • 收稿日期:2025-05-12 出版日期:2026-07-26 发布日期:2026-07-26
  • 通讯作者: 翟羽佳 E-mail:zhaiyujia_wish@163.com

3D Object Detection Model ICasA Based on Feature Enhancement

He Huaiqing, Zhai Yujia, Liu Haohan, Hui Kanghua   

  1. College of Computer and Artificial Intelligence, Civil Aviation University of China, Tianjin 300300, China
  • Received:2025-05-12 Online:2026-07-26 Published:2026-07-26

摘要: 针对3D目标检测模型中3D空间特征提取不充分的问题, 提出一种改进的ICasA模型. 首先, 引入焦点稀疏卷积, 对CasA模型的3D骨干网络进行改进, 增强对非空体素特征位置之间信息流的提取, 以获取更丰富的空间特征. 其次, 构建多尺度特征注意力融合模块, 嵌入2D骨干网络, 增强模型对多尺度特征的处理能力: 采用不同步长卷积加深2D骨干网络结构, 提升模型对全局特征的提取能力; 改进注意力融合方式, 在保留原始特征图的基础上实现对不同尺度特征不同位置的动态激活. 在数据集KITTI上的实验结果表明, 与CasA模型相比, ICasA模型对汽车、 行人和骑行者的3D平均检测精度AP@R40分别提高0.28百分点、 2.87百分点、 1.03百分点, 与现有先进模型相比, ICasA模型各类别目标检测效果更稳定, 有助于提升3D目标检测的精度. 

关键词: 3D目标检测, 焦点稀疏卷积, 多尺度特征注意力融合, 特征增强

Abstract: Aiming at the problem of insufficient extraction of 3D spatial features in 3D object detection models, we proposed an improved ICasA model. Firstly, we introduced focal sparse convolution to improve the 3D backbone network of CasA model, strengthening the extraction of information flow among non-empty voxel feature locations to obtain richer spatial features. Secondly, we constructed a multi-scale feature attention fusion module and embedded it into the 2D backbone network to enhance the model’s ability to handle multi-scale features. We adopted convolutions with different strides to deepen the structure of the 2D backbone network, improve the model’s ability to extract global features. We improved the attention fusion method to dynamically activate features of different scales and locations while preserving the original feature maps. The experimental results on the KITTI dataset show that compared with the CasA model, the ICasA model achieves 0.28 percentage points, 2.87 percentage points, and 1.03 percentage points increases in 3D average detection accuracy (AP@R40) for cars, pedestrians, and cyclists, respectively. Compared with existing advanced models, the ICasA model has more stable detection effect for all categories of objects, which helps to enhance the accuracy of 3D object detection.

Key words: 3D object detection, focal sparse convolution, multi-scale feature attention fusion, feature enhancement

中图分类号: 

  • TP391.4