Journal of Jilin University(Engineering and Technology Edition) ›› 2026, Vol. 56 ›› Issue (3): 819-829.doi: 10.13229/j.cnki.jdxbgxb.20240926

Previous Articles    

Efficient gating and target region attention based real-time video super-resolution

Le-ping LIN1,2(),Zhi SU2,Ning OUYANG1,2()   

  1. 1.Guangxi Key Laboratory of Wireless Broadband Communication and Signal Processing,Guilin University of Electronic Technology,Guilin 541004,China
    2.School of Information and Communication,Guilin University of Electronic Technology,Guilin 541004,China
  • Received:2024-08-23 Online:2026-03-01 Published:2026-03-31
  • Contact: Ning OUYANG E-mail:linleping@guet.edu.cn;ynou@guet.edu.cn

Abstract:

Facing complex video scenes with large motion amplitude, real-time video super-resolution algorithms are difficult to reconstruct texture details and occluded regions, based on generative adversarial network, a real-time video super-resolution method based on efficient gating and target region attention is proposed. The method firstly uses an efficient gating reconstruction network as a generative network to maintain high efficiency while adaptively selecting complex region information through a simplified gating mechanism to enhance the reconstruction results. Further, the method proposes a target region attention discriminative network to provide multiscale and spatio-temporal information feedback for the generative network, and acquires multiscale information of the complex video through the multiscale mechanism and ReLU linear attention; through the significant spatio-temporal discriminative module, it restricts the discriminative network to focus on the complex regions, and better acquires the spatio-temporal information of the complex regions. The experimental results show that the proposed method exhibits significant superiority over other SOTA algorithms, the model efficiency achieves an inference delay of 13.36 ms and a real-time score of 65.806, which indicates the efficient real-time performance of the model.

Key words: video super-resolution, real-time, generative adversarial network, gating mechanisms, target region attention

CLC Number: 

  • TP391.4

Fig.1

Illustration of the overall framework of generative network"

Fig.2

Illustration of the overall framework of NGM"

Fig.3

Illustration of RRB"

Fig.4

Illustration of the overall framework of TRADN"

Table 1

Experiment environment configuration"

名称环境配置
操作系统Ubuntu20.04
处理器Intel(R) Xeon(R) Gold6248 CPU @ 2.50 GHz
显卡Tesla A100
CUDA版本11.6
深度学习框架PyTorch 1.13.0
Python版本3.9
显存80 GB

Table 2

Quantitative comparison of different methods on three VSR testing sets"

方法REDS4Vid4UDM10
PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓
VESPCN22.290.594 40.334 219.700.500 20.574 726.220.791 60.143 6
SOFVSR24.310.660 20.317 221.070.530 30.356 827.630.816 90.131 1
TecoGAN22.990.630 30.291 624.810.770 70.200 931.970.900 90.088 8
FRVSR24.430.684 90.333 825.550.798 50.280 533.950.925 60.099 5
EGVSR29.750.838 20.109 025.830.796 30.137 636.960.941 10.067 6
STDO29.340.844 10.189 826.300.790 20.267 638.230.958 90.084 0
SSL29.810.835 60.223 726.610.819 60.245 338.260.957 20.079 6
EGTRA29.890.839 90.085 126.750.822 00.134 838.690.966 10.050 3

Fig.5

Visual comparisons on REDS4 testing set for 4× upscaling(Zoom in to see better visualization)"

Fig.6

Visual comparisons on Vid4 testing set for 4× upscaling(Zoom in to see better visualization)"

Fig.7

Visual comparisons on UDM10 testing set for 4× upscaling(Zoom in to see better visualization)"

Table 3

Comparison of model efficiency on the REDS4 testing set"

方法模型参数FLOPs推理延迟/ms实时推理REDS4评分REDS4 PSNR
VESPCN0.879M96.56G20.620.00122.29
SOFVSR1.640M226.12G75.13×0.00124.31
TecoGAN2.589M190.81G32.100.00122.99
FRVSR2.589M190.81G32.090.01424.43
EGVSR2.681M102.89G14.2550.81229.75
STDO4.843M2.76G35.7111.48529.34
SSL0.589M63.9G18.2543.11729.81
EGTRA2.177M60.69G13.3665.80629.89

Table 4

Validation of SCA module on three VSR testing sets"

方法

模型

参数

FLOPs推理延迟/msREDS4UDM10Vid4
PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓
SCA-o2.175M60.69G13.3229.770.835 00.085 326.680.81840.135 138.400.963 50.051 7
SCA-w2.177M60.69G13.3629.890.839 90.085 126.750.82200.134 838.690.966 10.050 3

Table 5

Validation of RRB module on three VSR testing sets"

方法

模型

参数

FLOPs推理延迟/msREDS4UDM10Vid4
PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓
RRB-o2.861M102.89G16.6829.940.841 20.085 026.780.824 50.134 838.760.966 40.050 3
RRB-w2.177M60.69G13.3629.890.839 90.085 126.750.822 00.134 838.690.966 10.050 3

Table 6

Validation of multiscale branches on three VSR testing sets"

方法

3×3

支路

5×5

支路

REDS4UDM10Vid4
PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓PSNR/dB↑SSIM↑LPIPS↓
Model l29.590.832 20.107 626.540.816 20.144 838.480.962 30.079 5
Model 229.680.834 10.095 026.610.816 30.142 938.540.963 30.058 3
Model 329.640.834 80.097 426.660.818 90.138 038.560.963 80.069 2
Model 429.890.839 90.085 126.750.822 00.134 838.690.966 10.050 3

Table 7

Validation of EGTRA modules on three VSR testing sets"

方法EGRNRLASSTDMREDS4UDM10Vid4

PSNR/

dB↑

SSIM↑LPIPS↓

PSNR/

dB↑

SSIM↑LPIPS↓

PSNR/

dB↑

SSIM↑LPIPS↓
SSTDM-o29.860.838 10.089 326.710.819 80.143 938.590.964 90.064 0
RLA-o29.590.832 20.107 626.540.816 20.144 838.480.964 10.079 5
EGRN-o29.410.830 80.085 226.290.814 20.135 138.200.960 50.050 8
Model29.890.839 90.085 126.750.822 00.134 838.690.966 10.050 3
[1] Claudio Rota, Marco Buzzelli, Simone Bianco, et al. Video restoration based on deep learning: a comprehensive survey[J]. Artificial Intelligence Review, 2023, 56(6): 5317-5364.
[2] Caballero J, Ledig C, Aitken A P, et al. Real-time video super-resolution with spatio-temporal networks and motion compensation[C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017: 4778-4787.
[3] Vemulapalli R, Brown M, Mehdi S M. Frame-recurrent video super-resolution[C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA,2018: 6626-6634.
[4] Wang L, Guo Y, Liu L, et al. Deep video super-resolution using HR optical flow estimation[J].IEEE Transactions on Image Processing, 2020, 29: 4323-4336.
[5] Chan K C K, Wang X, Yu K, et al. Basicvsr: the search for essential components in video super-resolution and beyond[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, TN, USA, 2021: 4947-4956.
[6] Chan K C K, Zhou S, Xu X,et al. Basicvsr++: improving video super-resolution with enhanced propagation and alignment[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA,2022: 5972-5981.
[7] Ouyang Ning, Zhi-shan Ou, Lin Le-ping. Video super-resolution network with gated high-low resolution frames[J]. Applied Sciences, 2023, 13(14): 1-16.
[8] Zhou X, Zhang L, Zhao X,et al. Video super-resolution transformer with masked inter & intra-frame attention[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024: 25399-25408.
[9] Cao Y, Wang C, Song C, et al. Real-time super-resolution system of 4k-video based on deep learning[C]∥2021 IEEE 32nd International Conference on Application-Specific Systems, Architectures and Processors(ASAP), NJ, USA, 2021: 69-76
[10] Xia Bin, He Jing-wen, Zhang Yu-lun, et al. Structured sparsity learning for efficient video super-resolution[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver, Canada. 2023: 22638-22647
[11] Xiao Jun, Jiang Xin-yang, Zheng Ning-xin, et al.Online video super-resolution with convolutional kernel bypass grafts[J]. IEEE Transactions on Multimedia, 2023, 25: 8972-8987.
[12] Li Gen, Ji Jie, Qin Ming-hai, et al. Towards high-quality and efficient video super-resolution via spatial-temporal data overfitting[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Vancouver, Canada. 2023: 10259-10269.
[13] Chu Meng-yu, Xie You, Laura Leal-Taixé, et al. Temporally coherent gans for video super-resolution (tecogan)[J]. arXiv Preprint, 2018, 1(2): 3.
[14] Chen Rui, Mu Yang, Zhang Yan. High-order relational generative adversarial network for video super-resolution[J]. Pattern Recognition, 2024, 146: 110059.
[15] Chen L, Chu X, Zhang X,et al. Simple baselines for image restoration[C]∥European Conference on Computer Vision. Cham: Springer, 2022: 17-33.
[16] Ding X, Zhang X, Han J, et al. Diverse branch block: Building a convolution as an inception-like unit[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, TN, USA, 2021: 10881-10890,.
[17] Cai H, Li J, Gan C, et al. Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris, France. 2023: 17302-17313.
[18] Ouyang D, He S, Zhan J, et al. Efficient multi-scale attention module with cross-spatial learning[C]∥ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). ArXiv, 2023, abs/2305.13563.
[19] Xue Tian-fan, Chen Bai-an, Wu Jia-jun, et al. Video enhancement with task-oriented flow[J]. International Journal of Computer Vision, 2019, 127: 1106-1125.
[20] Liu Ce, Sun De-qing. On Bayesian adaptive video super resolution[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2013, 36(2): 346-360.
[21] Yi P, Wang Z, Jiang K,et al. Progressive fusion video super-resolution network via exploiting non-local spatio-temporal correlations[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision,Seoul, Korea (South), 2019: 3106-3115.
[22] Nah S, Baik S, Hong S,et al. Ntire 2019 challenge on video deblurring and super-resolution: Dataset and study[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Long B, CA, USA, 2019, 1996-2005.
[23] Ignatov A, Romero A, Kim H,et al. Real-time video super-resolution on smartphones with deep learning, Mobile AI2021 challenge: Report[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 2021: 2535-2544.
[1] Liu ZHANG,Jia-xin LIANG,He LIU,Gui-xiang ZHANG,Wei LIU,Yan LI,Jia-bao ZHANG. Region of interest star map compression method based on real-time star location [J]. Journal of Jilin University(Engineering and Technology Edition), 2026, 56(3): 802-810.
[2] Lin LIN,Yu-xin CHEN,Wei-zhi NAI. Adaptive real⁃time gesture classification algorithm based on gesture frame sequence extraction [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(9): 3042-3048.
[3] Yan PIAO,Ji-yuan KANG. RAUGAN:infrared image colorization method based on cycle generative adversarial networks [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(8): 2722-2731.
[4] Hong XIAO,Xian-de LIU. Real-time acquisition and dynamic analysis of learning state based on hybrid intelligence [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(7): 2402-2408.
[5] Jing-hua ZHAO,Da LIU,Yu-qi ZHOU,Long WEN,Qian-yu LIU,Jie LIU,Fang-xi XIE. Air⁃fuel ratio control of engines based on Gaussian process regression intake prediction [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(6): 1854-1861.
[6] Guang-wen LIU,Qi-ying ZHAO,Chao WANG,Lian-yu Gao,Hua CAI,Qiang FU. Progressive recursive generative adversarial network-based single-image rain removal algorithm [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(4): 1363-1373.
[7] Yuan JI,Ya-qi YU. Optimization algorithm for speech facial video generation based on dense convolutional generative adversarial networks and keyframes [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(3): 986-992.
[8] Li-min YAN,Wei-ye JIN. Fast algorithm based on block direction prediction for AV1 intra encoding [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(3): 993-1000.
[9] Hong ZHAO,Yu-xuan MA,Fu-rong SONG. Image adversarial examples generation based on Diff⁃AdvGAN [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(12): 4052-4062.
[10] Xiang-long LUO,Xin-yu WEI,Mao-jun ZHAO,Ruo-chen LIU. Image dehazing algorithm based on contrast learning and generative adversarial network [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(10): 3296-3308.
[11] Xiao-yue WEN,Guo-min QIAN,Hua-hua KONG,Yue-jie MIU,Dian-hai WANG. TrafficPro: a framework to predict link speeds on signalized urban traffic network [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(8): 2214-2222.
[12] Xin-gang GUO,Ying-chen HE,Chao CHENG. Noise-resistant multistep image super resolution network [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(7): 2063-2071.
[13] Yu-kun ZHENG,Ru-yue SUN,Feng-ming LI,Yi-xiang LIU,Dong-guang LI,Rui SONG. Research and application of the centralized drive and control system for a hydraulic manipulator [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(11): 3358-3371.
[14] Feng-feng ZHOU,Tao YU,Yu-si FAN. Generative adversarial autoencoder integrated voting algorithm based on mass spectral data [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(10): 2969-2977.
[15] Nan ZHANG,Jian-hua SHI,Ji YI,Ping WANG. Real⁃time tracking method of underground moving target based on weighted centroid positioning [J]. Journal of Jilin University(Engineering and Technology Edition), 2023, 53(5): 1458-1464.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!