首页 /研究 /Visual Image Caption Generation for Service Robotics and Industrial Applications
OTHER

Visual Image Caption Generation for Service Robotics and Industrial Applications

Ren C. Luo, Yu‐Ting Hsu, Yu-Cheng Wen, Huan-Jun Ye

发表年份
2019
引用次数
38

摘要

Image caption generation is a task that generates a sentence from a raw image, which is mimicking the intelligence of human that can acquire knowledge from the view. The difficulty of this task is the combination of multimodal knowledge learning, i.e. recognition of objects, actions, scenes, human, etc. In order to perform semantic understanding for service robotics or other industrial applications, the caption must be enhanced for recognition of the objects in the confined environment. We propose a template-based augmentation method for improving the capability of object recognition while retaining the other capability of the image caption model. This work opens a new era of image caption generation training procedure that the caption dataset and the classification dataset can be combined to train the deep captioning model. We show in our experiments that our improved model outperforms the original model in SPICE metrics by 4 times.

关键词

Closed captioningComputer scienceArtificial intelligenceTask (project management)RoboticsService (business)SentenceImage (mathematics)Object (grammar)Computer vision

相关论文

查看 OTHER 分类全部论文