吉林大学学报(工学版) ›› 2026, Vol. 56 ›› Issue (3): 819-829.doi: 10.13229/j.cnki.jdxbgxb.20240926

• 计算机科学与技术 • 上一篇    

基于高效门控与目标区域关注的实时视频超分辨率

林乐平1,2(),苏治2,欧阳宁1,2()   

  1. 1.桂林电子科技大学 广西无线宽带通信与信号处理重点实验室,广西 桂林 541004
    2.桂林电子科技大学 信息与通信学院,广西 桂林 541004
  • 收稿日期:2024-08-23 出版日期:2026-03-01 发布日期:2026-03-31
  • 通讯作者: 欧阳宁 E-mail:linleping@guet.edu.cn;ynou@guet.edu.cn
  • 作者简介:林乐平(1980-),女,教授,博士.研究方向:机器学习,图像信号处理. E-mail:linleping@guet.edu.cn
  • 基金资助:
    国家自然科学基金项目(62001133);广西科技基地和人才专项项目(桂科AD19110060);广西自然科学基金项目(2017GXNSFBA198212);广西无线宽带通信与信号处理重点实验室基金项目(GXKL06200114)

Efficient gating and target region attention based real-time video super-resolution

Le-ping LIN1,2(),Zhi SU2,Ning OUYANG1,2()   

  1. 1.Guangxi Key Laboratory of Wireless Broadband Communication and Signal Processing,Guilin University of Electronic Technology,Guilin 541004,China
    2.School of Information and Communication,Guilin University of Electronic Technology,Guilin 541004,China
  • Received:2024-08-23 Online:2026-03-01 Published:2026-03-31
  • Contact: Ning OUYANG E-mail:linleping@guet.edu.cn;ynou@guet.edu.cn

摘要:

面对大运动幅度的复杂视频场景,实时视频超分辨率算法难以重建纹理细节、遮挡区域。本文基于生成对抗网络,提出了一种基于高效门控与目标区域关注的实时视频超分辨率方法。该方法首先使用高效门控重建网络作为生成网络,在保持高效的同时,通过简化的门控机制自适应选择复杂区域信息,以提升重建结果。进一步地,该方法提出了目标区域关注鉴别网络,为生成网络提供多尺度及时空信息反馈,通过多尺度机制和ReLU线性注意力获取复杂视频的多尺度信息;通过显著时空鉴别模块,限制鉴别网络关注复杂区域,以更好地获取复杂区域的时空信息。实验结果表明,所提方法相较于其他先进算法具有显著的优越性;在模型效率方面,实现了13.36 ms的推理延迟及65.806的实时得分,显示了模型高效的实时性能。

关键词: 视频超分辨率, 实时, 生成对抗网络, 门控机制, 目标区域关注

Abstract:

Facing complex video scenes with large motion amplitude, real-time video super-resolution algorithms are difficult to reconstruct texture details and occluded regions, based on generative adversarial network, a real-time video super-resolution method based on efficient gating and target region attention is proposed. The method firstly uses an efficient gating reconstruction network as a generative network to maintain high efficiency while adaptively selecting complex region information through a simplified gating mechanism to enhance the reconstruction results. Further, the method proposes a target region attention discriminative network to provide multiscale and spatio-temporal information feedback for the generative network, and acquires multiscale information of the complex video through the multiscale mechanism and ReLU linear attention; through the significant spatio-temporal discriminative module, it restricts the discriminative network to focus on the complex regions, and better acquires the spatio-temporal information of the complex regions. The experimental results show that the proposed method exhibits significant superiority over other SOTA algorithms, the model efficiency achieves an inference delay of 13.36 ms and a real-time score of 65.806, which indicates the efficient real-time performance of the model.

Key words: video super-resolution, real-time, generative adversarial network, gating mechanisms, target region attention

中图分类号: 

  • TP391.4

图1

生成网络的整体框架示意图"

图2

NGM整体框架示意图"

图3

RRB示意图"

图4

TRADN的整体框架示意图"

表1

实验环境配置"

名称环境配置
操作系统Ubuntu20.04
处理器Intel(R) Xeon(R) Gold6248 CPU @ 2.50 GHz
显卡Tesla A100
CUDA版本11.6
深度学习框架PyTorch 1.13.0
Python版本3.9
显存80 GB

表2

不同方法在3个VSR测试集上的定量比较"

方法REDS4Vid4UDM10
PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓
VESPCN22.290.594 40.334 219.700.500 20.574 726.220.791 60.143 6
SOFVSR24.310.660 20.317 221.070.530 30.356 827.630.816 90.131 1
TecoGAN22.990.630 30.291 624.810.770 70.200 931.970.900 90.088 8
FRVSR24.430.684 90.333 825.550.798 50.280 533.950.925 60.099 5
EGVSR29.750.838 20.109 025.830.796 30.137 636.960.941 10.067 6
STDO29.340.844 10.189 826.300.790 20.267 638.230.958 90.084 0
SSL29.810.835 60.223 726.610.819 60.245 338.260.957 20.079 6
EGTRA29.890.839 90.085 126.750.822 00.134 838.690.966 10.050 3

图5

在REDS4测试集上进行4倍放大的视觉比较(放大可查看更佳可视化效果)"

图 6

在Vid4测试集上进行4倍放大的视觉比较(放大可查看更佳可视化效果)"

图 7

在UDM10测试集上进行4倍放大的视觉比较(放大可查看更佳可视化效果)"

表 3

在REDS4测试集上对模型效率的比较"

方法模型参数FLOPs推理延迟/ms实时推理REDS4评分REDS4 PSNR
VESPCN0.879M96.56G20.620.00122.29
SOFVSR1.640M226.12G75.13×0.00124.31
TecoGAN2.589M190.81G32.100.00122.99
FRVSR2.589M190.81G32.090.01424.43
EGVSR2.681M102.89G14.2550.81229.75
STDO4.843M2.76G35.7111.48529.34
SSL0.589M63.9G18.2543.11729.81
EGTRA2.177M60.69G13.3665.80629.89

表4

三个VSR测试集上SCA模块的有效性验证"

方法

模型

参数

FLOPs推理延迟/msREDS4UDM10Vid4
PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓
SCA-o2.175M60.69G13.3229.770.835 00.085 326.680.81840.135 138.400.963 50.051 7
SCA-w2.177M60.69G13.3629.890.839 90.085 126.750.82200.134 838.690.966 10.050 3

表5

三个VSR测试集上RRB模块的有效性验证"

方法

模型

参数

FLOPs推理延迟/msREDS4UDM10Vid4
PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓
RRB-o2.861M102.89G16.6829.940.841 20.085 026.780.824 50.134 838.760.966 40.050 3
RRB-w2.177M60.69G13.3629.890.839 90.085 126.750.822 00.134 838.690.966 10.050 3

表6

三个VSR测试集上多尺度支路的有效性验证"

方法

3×3

支路

5×5

支路

REDS4UDM10Vid4
PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓
Model l29.590.832 20.107 626.540.816 20.144 838.480.962 30.079 5
Model 229.680.834 10.095 026.610.816 30.142 938.540.963 30.058 3
Model 329.640.834 80.097 426.660.818 90.138 038.560.963 80.069 2
Model 429.890.839 90.085 126.750.822 00.134 838.690.966 10.050 3

表7

三个VSR测试集上EGTRA各模块的有效性验证"

方法EGRNRLASSTDMREDS4UDM10Vid4

PSNR/

dB↑

SSIM↑LPIPS↓

PSNR/

dB↑

SSIM↑LPIPS↓

PSNR/

dB↑

SSIM↑LPIPS↓
SSTDM-o29.860.838 10.089 326.710.819 80.143 938.590.964 90.064 0
RLA-o29.590.832 20.107 626.540.816 20.144 838.480.964 10.079 5
EGRN-o29.410.830 80.085 226.290.814 20.135 138.200.960 50.050 8
Model29.890.839 90.085 126.750.822 00.134 838.690.966 10.050 3
[1] Claudio Rota, Marco Buzzelli, Simone Bianco, et al. Video restoration based on deep learning: a comprehensive survey[J]. Artificial Intelligence Review, 2023, 56(6): 5317-5364.
[2] Caballero J, Ledig C, Aitken A P, et al. Real-time video super-resolution with spatio-temporal networks and motion compensation[C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017: 4778-4787.
[3] Vemulapalli R, Brown M, Mehdi S M. Frame-recurrent video super-resolution[C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA,2018: 6626-6634.
[4] Wang L, Guo Y, Liu L, et al. Deep video super-resolution using HR optical flow estimation[J].IEEE Transactions on Image Processing, 2020, 29: 4323-4336.
[5] Chan K C K, Wang X, Yu K, et al. Basicvsr: the search for essential components in video super-resolution and beyond[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, TN, USA, 2021: 4947-4956.
[6] Chan K C K, Zhou S, Xu X,et al. Basicvsr++: improving video super-resolution with enhanced propagation and alignment[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA,2022: 5972-5981.
[7] Ouyang Ning, Zhi-shan Ou, Lin Le-ping. Video super-resolution network with gated high-low resolution frames[J]. Applied Sciences, 2023, 13(14): 1-16.
[8] Zhou X, Zhang L, Zhao X,et al. Video super-resolution transformer with masked inter & intra-frame attention[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024: 25399-25408.
[9] Cao Y, Wang C, Song C, et al. Real-time super-resolution system of 4k-video based on deep learning[C]∥2021 IEEE 32nd International Conference on Application-Specific Systems, Architectures and Processors(ASAP), NJ, USA, 2021: 69-76
[10] Xia Bin, He Jing-wen, Zhang Yu-lun, et al. Structured sparsity learning for efficient video super-resolution[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver, Canada. 2023: 22638-22647
[11] Xiao Jun, Jiang Xin-yang, Zheng Ning-xin, et al.Online video super-resolution with convolutional kernel bypass grafts[J]. IEEE Transactions on Multimedia, 2023, 25: 8972-8987.
[12] Li Gen, Ji Jie, Qin Ming-hai, et al. Towards high-quality and efficient video super-resolution via spatial-temporal data overfitting[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Vancouver, Canada. 2023: 10259-10269.
[13] Chu Meng-yu, Xie You, Laura Leal-Taixé, et al. Temporally coherent gans for video super-resolution (tecogan)[J]. arXiv Preprint, 2018, 1(2): 3.
[14] Chen Rui, Mu Yang, Zhang Yan. High-order relational generative adversarial network for video super-resolution[J]. Pattern Recognition, 2024, 146: 110059.
[15] Chen L, Chu X, Zhang X,et al. Simple baselines for image restoration[C]∥European Conference on Computer Vision. Cham: Springer, 2022: 17-33.
[16] Ding X, Zhang X, Han J, et al. Diverse branch block: Building a convolution as an inception-like unit[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, TN, USA, 2021: 10881-10890,.
[17] Cai H, Li J, Gan C, et al. Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris, France. 2023: 17302-17313.
[18] Ouyang D, He S, Zhan J, et al. Efficient multi-scale attention module with cross-spatial learning[C]∥ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). ArXiv, 2023, abs/2305.13563.
[19] Xue Tian-fan, Chen Bai-an, Wu Jia-jun, et al. Video enhancement with task-oriented flow[J]. International Journal of Computer Vision, 2019, 127: 1106-1125.
[20] Liu Ce, Sun De-qing. On Bayesian adaptive video super resolution[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2013, 36(2): 346-360.
[21] Yi P, Wang Z, Jiang K,et al. Progressive fusion video super-resolution network via exploiting non-local spatio-temporal correlations[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision,Seoul, Korea (South), 2019: 3106-3115.
[22] Nah S, Baik S, Hong S,et al. Ntire 2019 challenge on video deblurring and super-resolution: Dataset and study[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Long B, CA, USA, 2019, 1996-2005.
[23] Ignatov A, Romero A, Kim H,et al. Real-time video super-resolution on smartphones with deep learning, Mobile AI2021 challenge: Report[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 2021: 2535-2544.
[1] 张刘,梁嘉新,刘赫,张贵祥,刘威,李岩,章家保. 基于实时星点定位的感兴趣区域星图压缩方法[J]. 吉林大学学报(工学版), 2026, 56(3): 802-810.
[2] 林琳,陈雨欣,佴威至. 基于手势帧序列提取的自适应实时手势分类算法[J]. 吉林大学学报(工学版), 2025, 55(9): 3042-3048.
[3] 朴燕,康继元. RAUGAN:基于循环生成对抗网络的红外图像彩色化方法[J]. 吉林大学学报(工学版), 2025, 55(8): 2722-2731.
[4] 肖红,刘显德. 基于混合智能的学习状态实时采集与动态分析方法[J]. 吉林大学学报(工学版), 2025, 55(7): 2402-2408.
[5] 文斌,彭顺,杨超,沈艳军,李辉. 多深度自适应融合去雾生成网络[J]. 吉林大学学报(工学版), 2025, 55(6): 2103-2113.
[6] 赵靖华,刘妲,周宇麒,闻龙,刘倩妤,刘捷,解方喜. 基于高斯过程回归进气量预测的空燃比控制[J]. 吉林大学学报(工学版), 2025, 55(6): 1854-1861.
[7] 刘广文,赵绮莹,王超,高连宇,才华,付强. 基于渐进递归的生成对抗单幅图像去雨算法[J]. 吉林大学学报(工学版), 2025, 55(4): 1363-1373.
[8] 季渊,虞雅淇. 基于密集卷积生成对抗网络与关键帧的说话人脸视频生成优化算法[J]. 吉林大学学报(工学版), 2025, 55(3): 986-992.
[9] 严利民,金炜烨. 基于块方向预测的AV1快速帧内编码算法[J]. 吉林大学学报(工学版), 2025, 55(3): 993-1000.
[10] 赵宏,马宇轩,宋馥荣. 基于Diff-AdvGAN的图像对抗样本生成方法[J]. 吉林大学学报(工学版), 2025, 55(12): 4052-4062.
[11] 罗向龙,魏欣语,赵茂军,刘若辰. 融合对比学习和生成对抗网络的图像去雾算法[J]. 吉林大学学报(工学版), 2025, 55(10): 3296-3308.
[12] 张曦,库少平. 基于生成对抗网络的人脸超分辨率重建方法[J]. 吉林大学学报(工学版), 2025, 55(1): 333-338.
[13] 温晓岳,钱国敏,孔桦桦,缪月洁,王殿海. TrafficPro:一种针对城市信控路网的路段速度预测框架[J]. 吉林大学学报(工学版), 2024, 54(8): 2214-2222.
[14] 赖丹晖,罗伟峰,袁旭东,邱子良. 复杂环境下多模态手势关键点特征提取算法[J]. 吉林大学学报(工学版), 2024, 54(8): 2288-2294.
[15] 郭昕刚,何颖晨,程超. 抗噪声的分步式图像超分辨率重构算法[J]. 吉林大学学报(工学版), 2024, 54(7): 2063-2071.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!