Home /Research /Visual Image Caption Generation for Service Robotics and Industrial Applications
OTHER

Visual Image Caption Generation for Service Robotics and Industrial Applications

Ren C. Luo, Yu‐Ting Hsu, Yu-Cheng Wen, Huan-Jun Ye

Year
2019
Citations
38

Abstract

Image caption generation is a task that generates a sentence from a raw image, which is mimicking the intelligence of human that can acquire knowledge from the view. The difficulty of this task is the combination of multimodal knowledge learning, i.e. recognition of objects, actions, scenes, human, etc. In order to perform semantic understanding for service robotics or other industrial applications, the caption must be enhanced for recognition of the objects in the confined environment. We propose a template-based augmentation method for improving the capability of object recognition while retaining the other capability of the image caption model. This work opens a new era of image caption generation training procedure that the caption dataset and the classification dataset can be combined to train the deep captioning model. We show in our experiments that our improved model outperforms the original model in SPICE metrics by 4 times.

Keywords

Closed captioningComputer scienceArtificial intelligenceTask (project management)RoboticsService (business)SentenceImage (mathematics)Object (grammar)Computer vision

Related papers

Browse all OTHER papers