吉林大学学报(工学版) ›› 2026, Vol. 56 ›› Issue (2): 516-522.doi: 10.13229/j.cnki.jdxbgxb.20240251

• 计算机科学与技术 • 上一篇    

基于增强对象学习和注意力网络的视频描述方法

蔡晓东(),龙顺宏,梁焜峻   

  1. 桂林电子科技大学 信息与通信学院,广西 桂林 541004
  • 收稿日期:2024-03-12 出版日期:2026-02-01 发布日期:2026-03-17
  • 作者简介:蔡晓东(1971-),男,教授,博士. 研究方向:数据挖掘.E-mail: caixiaodong@guet.edu.cn
  • 基金资助:
    广西创新驱动发展专项项目(AA20302001)

Video captioning method based on enhanced object learning and attention networks

Xiao-dong CAI(),Shun-hong LONG,Kun-jun LIANG   

  1. School of Information and Communication,Guilin University of Electronic Technology,Guilin 541004,China
  • Received:2024-03-12 Online:2026-02-01 Published:2026-03-17

摘要:

在视频描述任务中,常见问题之一是对象描述不够具体,主要原因在于模型没有充分学习视频中的对象信息。同时,视频包含了丰富的特征信息,如对象信息、运动信息和上下文信息,这使得如何提升模型在生成描述时学习关键信息的能力成为一项具有挑战性的任务。为解决上述问题,本文提出了一种基于增强对象学习和注意力网络的方法。首先,设计了一种新的增强对象学习模块,旨在充分学习视频中的对象信息,从而实现对视频内容的准确描述;其次,构建了一种注意力网络,致力于有效关注不同类型的信息,以提升模型在生成描述时学习关键信息的能力。在MSVD和MSR-VTT数据集上的实验中,本文方法生成的描述展现出更高的具体性和准确性,同时在各项评价指标上均超过了目前的先进方法,有效验证了该方法的可行性。

关键词: 深度学习, 视频描述, 增强对象学习, 注意力网络

Abstract:

In video captioning tasks, one of the common problems is that the object caption is not specific enough, mainly because the model does not fully learn the information of the objects in the video. Meanwhile, videos contain abundant feature information, such as object information, motion information, and contextual information, making it a challenging task to enhance the model’s ability to learn key information when generating captions. To address the aforementioned problems, this paper proposes a method based on enhanced object learning and attention networks. Firstly, a new enhanced object learning module was designed to fully learn object information in videos, thereby achieving accurate caption of video content. Secondly, an attention network was constructed to dynamically adjust the weights of different types of information, thereby enhancing the model’s ability to learn key information when generating captions. In the experiments on the MSVD and MSR-VTT datasets, the caption generated by the method proposed in this paper showed a higher level of specificity and accuracy, and exceeded the current advanced methods in various evaluation indicators, effectively verifying the feasibility of the method.

Key words: deep learning, video captioning, enhanced object learning, attention network

中图分类号: 

  • TP391

图1

EOLM-AN模型的整体框架"

图2

EOLM结构图"

图3

AN结构图"

表1

在MSVD和MSR-VTT数据集上与最先进的方法进行比较 (%)"

方法MSVDMSR-VTT
B4MCRB4MCR
DMRM351.133.674.8
ORG-TRL554.336.495.273.943.628.850.962.1
POS-CG1452.534.188.771.342.028.248.761.6
LSRT1555.637.198.573.542.628.349.561.0
MA-LSTM1652.333.670.436.526.541.059.8
VADD1751.534.872.191.542.428.249.761.7
SwinBERT958.241.3120.677.541.929.953.862.1
TextKG1060.838.5105.275.143.729.652.462.4
MAN1160.137.1101.974.642.528.650.462.2
ViT/L141260.141.4121.578.244.430.357.263.4
HMN859.237.7104.075.143.529.051.562.7
EOLM-AN61.539.0106.777.544.229.452.163.6

表2

在MSVD和MSR-VTT数据集上的消融实验 (%)"

方法MSVDMSR-VTT
B4MCRB4MCR
EOLM61.338.1105.377.244.129.151.763.4
AN60.538.5106.276.243.729.351.962.9
EOLM-AN61.539.0106.777.544.229.452.163.6

图4

EOLM-AN模型在MSVD和MSR-VTT数据集上的可视化实例"

[1] Zhang J, Peng Y. Video captioning with object-aware spatio-temporal correlation and aggregation[J]. IEEE Transactions on Image Processing, 2020, 29: 6209-6222.
[2] Zanfir M, Marinoiu E, Sminchisescu C. Spatio-temporal attention models for grounded video captioning[C]∥Computer Vision-ACCV 2016: 13th Asian Conference on Computer Vision, Taipei, Taiwan, 2017: 104-119.
[3] Yang Z, Han Y, Wang Z. Catching the temporal regions-of-interest for video captioning[C]∥Proceedings of the 25th ACM International Conference on Multimedia, Mountain View, USA, 2017: 146-153.
[4] Zhang W, Wang X E, Tang S, et al. Relational graph learning for grounded video description generation[C]∥Proceedings of the 28th ACM International Conference on Multimedia, Seattle, USA, 2020: 3807-3828.
[5] Zhang Z, Shi Y, Yuan C, et al. Object relational graph with teacher-recommended learning for video captioning[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020: 13278-13288.
[6] Kanani C S, Saha S, Bhattacharyya P. Global object proposals for improving multi-sentence video descriptions[C]∥International Joint Conference on Neural Network, Montreal, Canada, 2021: 1-7.
[7] Parisotto E, Song F, Rae J, et al. Stabilizing transformers for reinforcement learning[C]∥International Conference on Machine Learning, Vienna, Austria, 2020: 7487-7498.
[8] Ye H, Li G, Qi Y, et al. Hierarchical modular network for video captioning[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, USA, 2022: 17939-17948.
[9] Lin K, Li L, Lin C C, et al. SwinBERT: end-to-end transformers with sparse attention for video captioning[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, USA, 2022: 17949-17958.
[10] Gu X, Chen G, Wang Y, et al. Text with knowledge graph augmented transformer for video captioning[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, USA, 2023: 18941-18951.
[11] Jing S, Zhang H, Zeng P, et al. Memory-based augmentation network for video captioning[J]. IEEE Transactions on Multimedia, 2023, 26: 2367-2379.
[12] Shen Y, Gu X, Xu K, et al. Accurate and fast compressed video captioning[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2023: 15558-15567.
[13] Wang J, Jiang W, Ma L, et al. Bidirectional attentive fusion with context gating for dense video captioning[C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 7190-7198.
[14] Wang B, Ma L, Zhang W, et al. Controllable video captioning with pos sequence guidance based on gated fusion network[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Korea, 2019: 2641-2650.
[15] Li L, Gao X, Deng J, et al. Long short-term relation transformer with global gating for video captioning[J]. IEEE Transactions on Image Processing, 2022, 31: 2726-2738.
[16] Xu J, Yao T, Zhang Y, et al. Learning multimodal attention LSTM networks for video captioning[C]∥Proceedings of the 25th ACM International Conference on Multimedia, Mountain View, USA, 2017: 537-545.
[17] Sun Z, Chen S, Zhong L. Visual-aware attention dual-stream decoder for video captioning[C]∥IEEE International Conference on Multimedia and Expo, Taipei, Taiwan, 2022: 1-6.
[1] 姚宗伟,陈辰,高振云,靳鸿鹏,荣浩,李学飞,黄虹溥,毕秋实. 基于合成图像数据集的挖掘机关键点识别[J]. 吉林大学学报(工学版), 2026, 56(1): 76-85.
[2] 王琳虹,刘宇阳,刘子昱,鹿应佳,张宇恒,黄桂树. 基于YOLOv5的轻量化桥梁缺陷识别[J]. 吉林大学学报(工学版), 2025, 55(9): 2958-2968.
[3] 廉敬,张继保,刘冀钊,张家骏,董子龙. 基于文本引导的人脸图像修复[J]. 吉林大学学报(工学版), 2025, 55(8): 2732-2740.
[4] 刘元宁,王星喆,黄子彧,张家晨,刘震. 基于多模态数据融合的胃癌患者生存预测模型[J]. 吉林大学学报(工学版), 2025, 55(8): 2693-2702.
[5] 袁靖舒,李武,赵兴雨,袁满. 基于BERTGAT-Contrastive的语义匹配模型[J]. 吉林大学学报(工学版), 2025, 55(7): 2383-2392.
[6] 徐慧智,郝东升,徐小婷,蒋时森. 基于深度学习的高速公路小目标检测算法[J]. 吉林大学学报(工学版), 2025, 55(6): 2003-2014.
[7] 张汝波,常世淇,张天一. 基于深度学习的图像信息隐藏方法综述[J]. 吉林大学学报(工学版), 2025, 55(5): 1497-1515.
[8] 李健,刘欢,李艳秋,王海瑞,关路,廖昌义. 基于THGS算法优化ResNet-18模型的图像识别[J]. 吉林大学学报(工学版), 2025, 55(5): 1629-1637.
[9] 文斌,丁弈夫,杨超,沈艳军,李辉. 基于自选择架构网络的交通标志分类算法[J]. 吉林大学学报(工学版), 2025, 55(5): 1705-1713.
[10] 李振江,万利,周世睿,陶楚青,魏巍. 基于时空Transformer网络的隧道交通运行风险动态辨识方法[J]. 吉林大学学报(工学版), 2025, 55(4): 1336-1345.
[11] 赵孟雪,车翔玖,徐欢,刘全乐. 基于先验知识优化的医学图像候选区域生成方法[J]. 吉林大学学报(工学版), 2025, 55(2): 722-730.
[12] 金虎,申玉生,方勇,于丽,周佳媚. 基于深度学习SSD算法的公路隧道衬砌细小裂缝识别[J]. 吉林大学学报(工学版), 2025, 55(11): 3653-3659.
[13] 蔡晓东,黄业洋,董丽芳. 基于增强正例与层间负例的语义相似性模型[J]. 吉林大学学报(工学版), 2025, 55(11): 3705-3714.
[14] 姜来为,王策,杨宏宇. 基于深度学习的多目标跟踪研究进展综述[J]. 吉林大学学报(工学版), 2025, 55(11): 3429-3445.
[15] 王威,孙钰洁,王新. 频率和空间特征融合的轻量级多尺度遥感图像场景分类网络[J]. 吉林大学学报(工学版), 2025, 55(10): 3361-3371.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!