Journal of Jilin University(Engineering and Technology Edition) ›› 2026, Vol. 56 ›› Issue (3): 662-669.doi: 10.13229/j.cnki.jdxbgxb.20240739

Previous Articles    

Autonomous grasping algorithm of robotic arm based on deep reinforcement learning

Cheng-jun TIAN(),Yu YAN,Ren-wei CUI,Jin-tong ZHANG   

  1. School of Electronic Information Engineering,Changchun University of Science and Technology,Changchun 130012,China
  • Received:2024-07-04 Online:2026-03-01 Published:2026-03-31

Abstract:

Aiming at the problems such as the lack of independent learning ability of traditional robotic arms and poor adaptability in unknown environments, this paper proposes a deep reinforcement learning algorithm combining prior knowledge, and modifies the design of reward function and experience pool to solve the problems such as low data quality, slow convergence speed and unsatisfactory learning effect in the early stage of training, while enhancing its generalization. The results on Pybullet simulation platform show that compared with the original algorithm, the convergence speed of the improved model is increased by 48.5%, and the success rate is increased by 8%. Compared with other mainstream algorithms, the convergence speed and success rate are significantly improved.

Key words: pattern recognition and intelligent system, deep reinforcement learning, priori knowledge, robot arm

CLC Number: 

  • TP183

Fig.1

Model design with prior knowledge"

Fig.2

Flow chart of multi-experience pool"

Fig.3

Six-degree-of-freedom robotic arm"

Fig.4

Three-finger flexible gripper"

Fig.5

Layout and training of the environment in Pybullet"

Fig.6

Comparison of reward values between original and improved algorithms"

Fig.7

Comparison of the success rate of the original algorithm and the improved algorithm"

Fig.8

Comparison of the convergence of the improved algorithm and the original algorithm"

Fig.9

Comparison of the success rate of the improved algorithm and the original algorithm"

Fig.10

Comparison of the convergence of the improved algorithm with other algorithms"

Fig.11

Comparison of the success rate of the improved algorithm with other algorithms"

Table 1

Multi-object grasping test"

实验物体每100次成功次数识别IOU
螺丝刀920.90
订书器990.94
香蕉800.88
锤子880.88
牙膏960.92
柠檬820.80
网球920.82
[1] 吕帅, 刘京. 基于深度强化学习的随机局部搜索启发式方法[J]. 吉林大学学报: 工学版, 2021, 51(4): 1420-1426.
Shuai Lyu, Liu Jing. Stochastic local search heuristic method based on deep reinforcement learning[J]. Journal of Jilin University (Engineering and Technology Edition), 2021, 51(4): 1420-1426.
[2] 李保罡, 王宇, 孔凡伟, 等. 基于智能反射表面辅助和信息年龄度量的安全状态更新[J]. 吉林大学学报:工学版, 2023, 53(10): 3014-3025.
Li Bao-gang, Wang Yu, Kong Fan-wei, et al. Security status updates based on intelligent reflecting surface assistance and age of information metrics[J]. Journal of Jilin University (Engineering and Technology Edition), 2023, 53(10): 3014-3025.
[3] Schaul T, Quan J, Antonoglou I, et al. Prioritized experience replay[J/OL]. arXiv preprint arXiv:, 2015.
[4] Rauber P, Ummadisingu A, Mutz F, et al. Hindsight policy gradients[J/OL]. arXiv preprint arXiv:, 2017.
[5] Haarnoja T, Zhou A, Hartikainen K, et al. Soft actor-critic algorithms and applications[J/OL]. arXiv preprint arXiv:, 2018.
[6] Rakelly K, Zhou A, Finn C, et al. Efficient off-policy meta-reinforcement learning via probabilistic context variables[C]∥International Conference on Machine Learning, PMLR, 2019: 5331-5340.
[7] 庄伟超, 丁昊楠, 董昊轩, 等. 信号交叉口网联电动汽车自适应学习生态驾驶策略[J]. 吉林大学学报: 工学版, 2023, 53(1): 82-93.
Zhuang Wei-chao, Ding Hao-nan, Dong Hao-xuan, et al. Learning based eco⁃driving strategy of connected electric vehicle at signalized intersection[J]. Journal of Jilin University (Engineering and Technology Edition), 2023, 53(1): 82-93.
[8] 高敬鹏, 王国轩, 高路. 基于异步合作更新的LSTM-MADDPG多智能体协同决策算法[J]. 吉林大学学报: 工学版, 2024, 54(3): 797-806.
Gao Jing-peng, Wang Guo-xuan, Gao Lu. LSTM⁃MADDPG multi⁃agent cooperative decision algorithm based on asynchronous collaborative update[J]. Journal of Jilin University (Engineering and Technology Edition), 2024, 54(3): 797-806.
[9] 张强, 文闻, 周晓东, 等. 基于改进TD3算法的机械臂智能规划方法研究[J]. 智能科学与技术学报, 2022, 4(2): 223-232.
Zhang Qiang, Wen Wen, Zhou Xiao-dong, et al. Research on the manipulator intelligent trajectory planning method based on the improved TD3 algorithm[J]. Chinese Journal of Intelligent Science and Technology, 2022, 4(2): 223-232.
[10] 鲜斌, 张诗婧, 韩晓薇, 等. 基于强化学习的无人机吊挂负载系统轨迹规划[J]. 吉林大学学报:工学版, 2021, 51(6): 2259-2267.
Xian Bin, Zhang Shi-jing, Han Xiao-wei, et al. Trajectory planning for unmanned aerial vehicle slung⁃payload aerial transportation system based on reinforcement learning[J]. Journal of Jilin University (Engineering and Technology Edition), 2021, 51(6): 2259-2267.
[11] 康朝海, 孙超, 荣垂霆, 等. 基于动态延迟策略更新的TD3 算法[J]. 吉林大学学报: 信息科学版, 2020, 38(4): 474-481.
Kang Chao-hai, Sun Chao, Rong Chui-ting, et al. TD3 Algorithm with dynamic delayed policy update[J]. Journal of Jilin University (Information Science Edition), 2020, 38(4): 474-481.
[12] 刘庆强, 刘鹏云. 基于优先级经验回放的SAC强化学习算法[J]. 吉林大学学报: 信息科学版, 2021, 39(2): 192-199.
Liu Qing-qiang, Liu Peng-yun. Soft actor critic reinforcement learning with prioritized experience replay[J]. Journal of Jilin University (Information Science Edition), 2021, 39(2): 192-199.
[13] 赵宏伟, 陈霄, 龙曼丽, 等. 基于改进PLSA分类器的目标分类算法[J]. 吉林大学学报: 工学版, 2012, 42(): 231-235.
Zhao Hong-wei, Chen Xiao, Long Man-li, et al. Object classification algorithm based on improved PLSA[J]. Journal of Jilin University (Engineering and Technology Edition), 2012, 42(Sup.1): 231-235.
[14] 刘勇, 徐雷, 张楚晗. 面向文本游戏的深度强化学习模型[J]. 吉林大学学报: 工学版, 2022, 52(3): 666-674.
Liu Yong, Xu Lei, Zhang Chu-han. Deep reinforcement learning model for text games[J]. Journal of Jilin University (Engineering and Technology Edition), 2022, 52(3): 666-674.
[1] Zi-hao SHEN,Yong-sheng GAO,Hui WANG,Pei-qian LIU,Kun LIU. Deep deterministic policy gradient caching method for privacy protection in Internet of Vehicles [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(5): 1638-1647.
[2] Wei-chao HU,Zhen-ming YANG,Peng-cheng YU,Yan-yan CHEN,She-qiang MA. Modeling interaction policy of autonomous vehicle and pedestrian based on deep reinforcement learning [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(10): 3180-3188.
[3] Sheng-jie ZHU,Xuan WANG,Fang XU,Jia-qi PENG,Yuan-chao WANG. Multi-scale normalized detection method for airborne wide-area remote sensing images [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(8): 2329-2337.
[4] Guang-he ZHU,Zhi-qiang ZHU,Yi-ping YUAN. Deep reinforcement learning optimization scheduling algorithm for continuous production line [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(7): 2086-2092.
[5] Jing-peng GAO,Guo-xuan WANG,Lu GAO. LSTM⁃MADDPG multi⁃agent cooperative decision algorithm based on asynchronous collaborative update [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(3): 797-806.
[6] Jian ZHANG,Qing-yang LI,Dan LI,Xia JIANG,Yan-hong LEI,Ya-ping JI. Merging guidance of exclusive lanes for connected and autonomous vehicles based on deep reinforcement learning [J]. Journal of Jilin University(Engineering and Technology Edition), 2023, 53(9): 2508-2518.
[7] Yan-tao TIAN,Yan-shi JI,Huan CHANG,Bo XIE. Deep reinforcement learning augmented decision⁃making model for intelligent driving vehicles [J]. Journal of Jilin University(Engineering and Technology Edition), 2023, 53(3): 682-692.
[8] Wei-chao ZHUANG,Hao-nan Ding,Hao-xuan DONG,Guo-dong YIN,Xi WANG,Chao-bin ZHOU,Li-wei XU. Learning based eco⁃driving strategy of connected electric vehicle at signalized intersection [J]. Journal of Jilin University(Engineering and Technology Edition), 2023, 53(1): 82-93.
[9] Yong-jie MA,Min CHEN. Dynamic multi⁃objective optimization algorithm based on Kalman filter prediction strategy [J]. Journal of Jilin University(Engineering and Technology Edition), 2022, 52(6): 1442-1458.
[10] Yong LIU,Lei XU,Chu-han ZHANG. Deep reinforcement learning model for text games [J]. Journal of Jilin University(Engineering and Technology Edition), 2022, 52(3): 666-674.
[11] Zhong-li WANG,Hao WANG,Yan SHEN,Bai-gen CAI. A driving decision⁃making approach based on multi⁃sensing and multi⁃constraints reward function [J]. Journal of Jilin University(Engineering and Technology Edition), 2022, 52(11): 2718-2727.
[12] Da-ke ZHOU,Chao ZHANG,Xin YANG. Self-supervised 3D face reconstruction based on multi-scale feature fusion and dual attention mechanism [J]. Journal of Jilin University(Engineering and Technology Edition), 2022, 52(10): 2428-2437.
[13] Ya-hui ZHAO,Fei-yang YANG,Zhen-guo ZHANG,Rong-yi CUI. Korean text structure discovery based on reinforcement learning and attention mechanism [J]. Journal of Jilin University(Engineering and Technology Edition), 2021, 51(4): 1387-1395.
[14] Shuai LYU,Jing LIU. Stochastic local search heuristic method based on deep reinforcement learning [J]. Journal of Jilin University(Engineering and Technology Edition), 2021, 51(4): 1420-1426.
[15] Zhuo-jun XU,Wen-ting YANG,Cheng-zhi YANG,Yan-tao TIAN,Xiao-jun WANG. Improved residual neural network algorithm for radar intra-pulse modulation classification [J]. Journal of Jilin University(Engineering and Technology Edition), 2021, 51(4): 1454-1460.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!