Shaping in reinforcement learning by knowledge transferred from human-demonstrations of a simple similar task

Wang, Guo-Fang; Fang, Zhou; Li, Ping

doi:10.3233/JIFS-17052

Shaping in reinforcement learning by knowledge transferred from human-demonstrations of a simple similar task

Article type: Research Article

Authors: Wang, Guo-Fang^a | Fang, Zhou^{a; *} | Li, Ping^b

Affiliations: [a] School of Aeronautics and Astronautics, Zhejiang University, Hangzhou, China | [b] School of Control Science and Engineering, Zhejiang University, Hangzhou, China

Correspondence: [*] Corresponding author. Zhou Fang, School of Aeronautics and Astronautics, Zhejiang University, Hangzhou 310027, China. E-mail: [email protected].

Abstract: Reusing knowledge obtained in other related but different tasks to accelerate the learning procedure of reinforcement learning (RL) has attracted more and more attention and expert knowledge transfer is the root cause of positive effect. Nevertheless, compared with acquiring knowledge by RL training in source tasks, this paper proposes to transfer knowledge contained in human-demonstrations of source tasks. Based on this, three specific forms of knowledge in total are mined from demonstration trajectories to be reused in the target task to shape RL and all of them are closely associated with the similarity between states of different tasks which can be measured by Euclidean distance via human-supplied inter-task mappings. In more detail, the similarity between the target state and the most similar state in source samples, the proportion of different actions among the k-NN of the target state in source samples and the proportion of different actions under a constant similarity with the target state in source samples are respectively selected to initialize the value of state-action function. Simulation experiments of mountain car problems with different difficulties and different dimensions suggest that all the three shaping methods could obviously speed up RL. In comparison, it can also be found that the two latter methods are more robust and efficient to the quality of human demonstrations as it takes more source samples’ information into consideration.

Keywords: Transfer, human-demonstrations, shaping, reinforcement learning

DOI: 10.3233/JIFS-17052

Journal: Journal of Intelligent & Fuzzy Systems, vol. 34, no. 1, pp. 711-720, 2018

Published: 12 January 2018

Price: EUR 27.50

North America

IOS Press, Inc.
6751 Tepper Drive
Clifton, VA 20124
USA

Tel: +1 703 830 6300
Fax: +1 703 830 2300
[email protected]

For editorial issues, like the status of your submitted paper or proposals, write to [email protected]

Europe

IOS Press
Nieuwe Hemweg 6B
1013 BG Amsterdam
The Netherlands

Tel: +31 20 688 3355
Fax: +31 20 687 0091
[email protected]

For editorial issues, permissions, book requests, submissions and proceedings, contact the Amsterdam office [email protected]

Asia

Inspirees International (China Office)
Ciyunsi Beili 207(CapitaLand), Bld 1, 7-901
100025, Beijing
China

Free service line: 400 661 8717
Fax: +86 10 8446 7947
[email protected]

For editorial issues, like the status of your submitted paper or proposals, write to [email protected]

如果您在出版方面需要帮助或有任何建, 件至: [email protected]

Share this:

North America

Europe

Asia