Journal of Jilin University(Engineering and Technology Edition) ›› 2026, Vol. 56 ›› Issue (9): 2444-2455.doi: 10.13229/j.cnki.jdxbgxb.20250237

Previous Articles    

Algorithm of coverage path planning for unmanned aerial vehicle based on improved DQN

Xing-wang WANG1(),Jin-kuan LUO2,Zi-yang ZHANG2,Yu-pu CHI2,Jin-you JIANG2,Mei-ming YU3()   

  1. 1.College of Computer Science and Technology,Jilin University,Changchun 130012,China
    2.College of Software,Jilin University,Changchun 130012,China
    3.Public Computer Education and Research Center,Jilin University,Changchun 130012,China
  • Received:2025-03-21 Online:2026-09-01 Published:2026-09-07
  • Contact: Mei-ming YU E-mail:xww@jlu.edu.cn;yumm@jlu.edu.cn

Abstract:

Aiming at the problems of low coverage, high path redundancy and slow convergence speed of traditional DQN in UAV coverage path planning, an improved DQN model based on multi-mechanism fusion was proposed. Firstly, the workspace of the UAV was modeled as a two-dimensional grid map. It was preprocessed and the reward function was designed. Secondly, based on the traditional DQN, three mechanisms of Double DQN, Dueling Network and Priority Experience Replay were gradually integrated, and a multi-scale convolutional neural network was built using techniques such as one-hot encoding to construct an improved DQN model. Finally, a visual simulation experiment was designed, and an ablation experiment was performed using evaluation indicators such as grid coverage. The experimental results show that the improved DQN model performs well in the coverage path planning task, and its performance is better than that of traditional DQN and other models. Compared with traditional DQN, its grid coverage rate is increased by 8.28%. Also, the grid repetition rate, path length, number of collisions with obstacles, and number of times exceeding boundaries are reduced by 28.44%, 45.14%, 77.31%, and 90.14%. The network convergence speed is significantly accelerated as well.

Key words: computer applications, deep reinforcement learning, coverage path planning, unmanned aerial vehicle, grid maps

CLC Number: 

  • TP391

Fig.1

Schematic diagram of grid maps construction"

Fig.2

Schematic diagram of obstacle regularization"

Fig.3

Schematic diagram of DQN architecture"

Fig.4

Schematic diagram of Dueling network"

Fig.5

Structure of SumTree"

Fig.6

Architecture of improved DQN"

Fig.7

Neural network architecture of Improved DQN"

Table 1

Parameter settings for simulation experiments"

参数含义
ω网络学习率1e-4
γ折扣因子0.99
?见式(9)1e-5
μ见式(9)0.6
σ见式(11)0.4
H二维栅格地图高度10
W二维栅格地图宽度10
frame_max最大训练帧数1 000 000
frame_start训练开始帧数10 000
|D|回放经验池最大样本容量100 000
|K|可选动作的总数4
Batch_size每次训练选取的样本数32
epsilon_max最大贪婪系数1.0
epsilon_min最小贪婪系数0.01
epsilon_decay至最小贪婪系数的帧数30 000
max_steps每回合最大步数300
agent_num智能体数量1
update_freq网络参数更新间隔/帧1 000
obstacles_per障碍物栅格的占比/%16

Table 2

Experimental results of parameter sensitivity analysis"

改进型DQN模型相关参数

栅格覆盖

率/%

栅格重复

率/%

路径长度碰撞障碍物次数超出边界次数
ω=1e-4、γ=0.99、G?i,j=G?(i',j')时奖励值+0.2597.9035.83154.415.460.34
ω=1e-398.3731.33144.913.970.51
ω=1e-583.8971.03291.7513.111.97
γ=0.896.9143.36173.433.402.44
γ=0.997.8539.97164.514.411.54
G?i,j=G?(i',j')时奖励值+092.4153.94230.0015.792.01
G?i,j=G?(i',j')时奖励值+0.598.0435.67158.679.030.27

Fig.8

Network training loss curve"

Fig.9

Network training reward curve"

Fig.10

Coverage path simulation diagram"

Table 3

Ablation experiment results"

DDQN

Dueing

network

PER栅格覆盖率/%栅格重复率/%

路径

长度

碰撞障碍物次数超出边界次数

训练

时间

#Param

/M

#FLOPs

/G

???89.6264.27281.4524.063.453:22:043.721.11
???93.9261.83273.6623.780.854:14:593.721.11
???93.7256.13233.7912.723.653:56:057.131.22
???94.4751.18215.9911.692.623:26:303.721.11
???94.8250.40219.9219.370.714:31:457.131.22
???96.2644.76187.6610.760.774:07:367.131.22
???95.9448.76201.1610.053.343:38:553.721.11
???97.9035.83154.415.460.344:36:027.131.22

Table 4

Ablation experiment results"

模型H,Wmax_steps栅格覆盖率/%栅格重复率/%路径长度碰撞障碍物次数超出边界次数
DQN1030089.6264.27281.4524.063.45
DQN+①1030093.9261.83273.6623.780.85
DQN+②1030093.7256.13233.7912.723.65
DQN+③1030094.4751.18215.9911.692.62
改进型DQN1030097.9035.83154.415.460.34
DQN1235084.6863.46349.7020.3626.14
DQN+①1235086.1464.84348.7613.0743.30
DQN+②1235085.1661.47345.8924.4531.46
DQN+③1235089.8058.61329.0818.7612.15
改进型DQN1235094.1357.00320.9513.005.46
DQN1540080.7753.10399.9829.4919.75
DQN+①1540084.6450.86400.0030.8430.76
DQN+②1540083.1451.68399.9930.5135.20
DQN+③1540087.6052.56399.9418.1725.78
改进型DQN1540092.7049.66391.9914.294.96
[1] Buchelt A, Adrowitzer A, Kieseberg P, et al. Exploring artificial intelligence for applications of drones in forest ecology and management[J]. Forest Ecology and Management, 2024, 551: No.121530.
[2] Bakirci M. Smart city air quality management through leveraging drones for precision monitoring[J]. Sustainable Cities and Society, 2024, 106: No.105390.
[3] Ibrahim Z T, He J. Slam technology on disaster response[J]. World Journal of Engineering and Technology, 2024, 12(3): 695-714.
[4] Xing B, Wang X, Yang L, et al. An algorithm of complete coverage path planning for unmanned surface vehicle based on reinforcement learning[J]. Journal of Marine Science and Engineering, 2023, 11(3): No.645.
[5] Kim J. Autonomous robot vacuum system composed of a cleaner robot and a dust storage robot[J]. Journal of the Franklin Institute, 2024, 361(11): No.106938.
[6] Wan S, Chen Z, Dong J. An efficiency-based interactive dynamic technique with interval-valued hesitant fuzzy constraint cone for rescue route planning[J]. Expert Systems with Applications, 2023, 231: No.120648.
[7] Dijkstra E W. A note on two problems in connexion with graphs[J]. Numerische Mathematik, 1959, 1(1): 269-271.
[8] Hart P E, Nilsson N J, Raphael B. A formal basis for the heuristic determination of minimum cost paths[J]. IEEE transactions on Systems Science and Cybernetics, 1968, 4(2): 100-107.
[9] Xu R, Xia Y, Chen P. A Multi-Objective Particle Swarm Optimization Algorithm for Drone Path Planning in Forest Firefighting[C]∥The 11th International Conference on Electrical and Electronics Engineering(ICEEE), Marmaris,Turkiye,2024:523-527.
[10] 唐颂, 吴建源. 基于改进遗传算法的协同航迹规划方法[J]. 电光与控制, 2024, 31(7): 8-12, 26.
Tang Song, Wu Jian-yuan. A Cooperative Trajectory Planning Method Based on Improved Genetic Algorithm[J]. Electronics Optics & Control, 2024, 31(7): 8-12, 26.
[11] Luo J, Shun H, Hao Y. A path planning method for comprehensive area coverage utilizing neuron activity[C]∥The 5th International Conference on Robotics, Intelligent Control and Artificial Intelligence(RICAI),Hangzhou,China, 2023: 379-383.
[12] 屈博琛. 基于深度学习自动微分的无人机路径规划[J]. 自动化与仪表, 2025, 40(2): 66-72.
Qu Bo-chen. Drone path planning based on deep learning with automatic differentiation[J]. Automation & Instrumentation, 2025, 40(2): 66-72.
[13] Mnih V, Kavukcuoglu K, Silver D, et al. Playing atari with deep reinforcement learning[J/OL].[2025-02-16]. .
[14] van Hasselt H, Guez A, Silver D. Deep reinforcement learning with double Q-learning[C]∥Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, Phoenix, Arizona, 2016, 30(1): 2094-2100.
[15] Wang Z, Schaul T, Hessel M, et al. Dueling network architectures for deep reinforcement learning[C]∥International Conference on Machine Learning, New York, USA, 2016: 1995-2003.
[16] Schaul T, Quan J, Antonoglou I, et al. Prioritized experience replay[J/OL]. [2025-01-22]. .
[17] Watkins C J C H. Learning from delayed rewards[D]. Cambridge: King's College, University of Cambridge,1989.
[18] Jaderberg M, Simonyan K, Zisserman A. Spatial transformer networks[J]. Advances in Neural Information Processing Systems, 2015, 28: 2017-2025.
[19] Hu J, Shen L, Sun G. Squeeze-and-excitation networks[C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018:7132-7141.
[20] Brockman G, Cheung V, Pettersson L, et al. Openai gym[J/OL].[2025-01-26]. .
[1] Tao JU,Wen-jin ZHANG,Yao YANG,Jiu-yuan HUO. A second order decision-making dynamic offloading method for vehicle edge computing tasks [J]. Journal of Jilin University(Engineering and Technology Edition), 2026, 56(7): 2006-2019.
[2] Cheng-jun TIAN,Yu YAN,Ren-wei CUI,Jin-tong ZHANG. Autonomous grasping algorithm of robotic arm based on deep reinforcement learning [J]. Journal of Jilin University(Engineering and Technology Edition), 2026, 56(3): 662-669.
[3] Qing-lin AI,Yuan-xiao LIU,Jia-hao YANG. Small target swmantic segmentation method based MFF-STDC network in complex outdoor environments [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(8): 2681-2692.
[4] Zi-hao SHEN,Yong-sheng GAO,Hui WANG,Pei-qian LIU,Kun LIU. Deep deterministic policy gradient caching method for privacy protection in Internet of Vehicles [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(5): 1638-1647.
[5] Tao XU,Shuai-di KONG,Cai-hua LIU,Shi LI. Overview of heterogeneous confidential computing [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(3): 755-770.
[6] Yao-ping ZENG,Yu-ting XIA,Shi-sen CHEN,Yue-qiang LIU,Wei-wei JIANG. Energyefficient offloading strategy research for multiUAV assisted communication [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(12): 4083-4092.
[7] Wei-chao HU,Zhen-ming YANG,Peng-cheng YU,Yan-yan CHEN,She-qiang MA. Modeling interaction policy of autonomous vehicle and pedestrian based on deep reinforcement learning [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(10): 3180-3188.
[8] Hai-yan HUANG,Hong-sheng ZHANG,Lin-lin LIANG,Chun-li WANG,Xue-jun ZHANG. Analysis of intelligent communication system with multi-UAVs based on SC/MRC [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(10): 3401-3409.
[9] Yi TANG,Yang PAN,Ming GAO,Hong-chen YI,An-qi WEI. Multi spectral image matching algorithm of unmanned aerial vehicle based on affine invariant operator [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(7): 2080-2085.
[10] Guang-he ZHU,Zhi-qiang ZHU,Yi-ping YUAN. Deep reinforcement learning optimization scheduling algorithm for continuous production line [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(7): 2086-2092.
[11] Bin XIAN,Guang-yi WANG,Jia-ming CAI. Nonlinear robust control design for multi unmanned aerial vehicles suspended payload transportation system [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(6): 1788-1795.
[12] Dian-wei WANG,Chi ZHANG,Jie FANG,Zhi-jie XU. UAV target tracking algorithm based on high resolution siamese network [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(5): 1426-1434.
[13] Jing-peng GAO,Guo-xuan WANG,Lu GAO. LSTM⁃MADDPG multi⁃agent cooperative decision algorithm based on asynchronous collaborative update [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(3): 797-806.
[14] Jian ZHANG,Qing-yang LI,Dan LI,Xia JIANG,Yan-hong LEI,Ya-ping JI. Merging guidance of exclusive lanes for connected and autonomous vehicles based on deep reinforcement learning [J]. Journal of Jilin University(Engineering and Technology Edition), 2023, 53(9): 2508-2518.
[15] Ze-qiang ZHANG,Wei LIANG,Meng-ke XIE,Hong-bin ZHENG. Elite differential evolution algorithm for mixed⁃model two⁃side disassembly line balancing problem [J]. Journal of Jilin University(Engineering and Technology Edition), 2023, 53(5): 1297-1304.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!