Affiliations: [a] Department of Computer Science, University of Bari Aldo Moro, Italy IT | [b] Interaction Lab, Heriot-Watt University, Edinburgh Centre for Robotics Edinburgh, Scotland, UK
Corresponding author: Pierpaolo Basile, Department of Computer Science, University of Bari Aldo Moro, Italy IT. E-mail: [email protected].
Abstract: In this paper, we propose a framework based on Hierarchical Reinforcement Learning for dialogue management in a Conversational Recommender System scenario. The framework splits the dialogue into more manageable tasks whose achievement corresponds to goals of the dialogue with the user. The framework consists of a meta-controller, which receives the user utterance and understands which goal should pursue, and a controller, which exploits a goal-specific representation to generate an answer composed by a sequence of tokens. The modules are trained using a two-stage strategy based on a preliminary Supervised Learning stage and a successive Reinforcement Learning stage.