Journal of Jilin University (Information Science Edition) ›› 2020, Vol. 38 ›› Issue (4): 474-481.
Previous Articles Next Articles
KANG Chaohai,SUN Chao,RONG Chuiting,LIU Pengyun
Received:
Online:
Published:
Abstract: In the field of deep reinforcement learning, in order to further reduce the impact of value overestimation on policy estimation in TD3 ( Twin Delayed Deep Deterministic Policy Gradients) and accelerate the efficiency of model learning,a DD-TD3 ( Twin Delayed Deep Deterministic Policy Gradients with Dynamic Delayed Policy Update) is proposed. The delay update step size of the actor network is guided by the dynamic difference between the latest loss of the critic network and its exponential weighted moving average. Experimental results show that compared to the original TD3 algorithm that obtain high reward value in the 2 000 steps,the DD-TD3 method can learn the optimal control strategy in about 1 000 steps and obtain a higher reward value, thereby the efficiency of finding the optimal strategy is improved.
Key words: deep reinforcement learning, twin delayed deep deterministic policy gradients ( TD3) , dynamic delayed policy update
CLC Number:
KANG Chaohai, SUN Chao, RONG Chuiting, LIU Pengyun. TD3 Algorithm with Dynamic Delayed Policy Update[J].Journal of Jilin University (Information Science Edition), 2020, 38(4): 474-481.
0 / / Recommend
Add to citation manager EndNote|Reference Manager|ProCite|BibTeX|RefWorks
URL: https://xuebao.jlu.edu.cn/xxb/EN/
https://xuebao.jlu.edu.cn/xxb/EN/Y2020/V38/I4/474
Cited