Affiliations: [a] Faculty of Computer Science and Engineering, Delhi Technological University, Shahbad Daulatpur, Delhi – 42, India | [b] Department of Information Technology, Delhi Technological University, Shahbad Daulatpur, Delhi – 42, India
Corresponding author: Akshi Kumar, Faculty of Computer Science Engineering, Delhi Technological University, Shahbad Daulatpur, Main Bawana Road, Delhi – 42, India. E-mail: firstname.lastname@example.org.
Note:  These authors have equal contribution.
Abstract: Automatic captioning of Images has been explored extensively in the past 10 to 15 years. It is one of the elementary problems in Computer Vision and Natural Language Processing and has vast array of applications in the real world. In this survey, we aim to study different approaches used for the generation of image captions in a chronological manner starting from the basic template based caption generation model to using Neural Networks combined with external world knowledge. We review existing models in detail, highlighting the involved methodologies and improvements in the same that have occurred in time. We gave an overview to the standard image datasets and the evaluation measures developed to discern the quality of generated image captions. Apart from the basic benchmarks we also note speed and accuracy improvements in all the different approaches. Finally, we investigate further possibilities in automatic image caption generation.
Keywords: Computer Vision, image captioning, deep learning object recognition, Natural Language Processing