Visual Image Caption Generation for Service Robotics and Industrial Applications
Ren C. Luo, Yu‐Ting Hsu, Yu-Cheng Wen, Huan-Jun Ye
- 发表年份
- 2019
- 引用次数
- 38
摘要
Image caption generation is a task that generates a sentence from a raw image, which is mimicking the intelligence of human that can acquire knowledge from the view. The difficulty of this task is the combination of multimodal knowledge learning, i.e. recognition of objects, actions, scenes, human, etc. In order to perform semantic understanding for service robotics or other industrial applications, the caption must be enhanced for recognition of the objects in the confined environment. We propose a template-based augmentation method for improving the capability of object recognition while retaining the other capability of the image caption model. This work opens a new era of image caption generation training procedure that the caption dataset and the classification dataset can be combined to train the deep captioning model. We show in our experiments that our improved model outperforms the original model in SPICE metrics by 4 times.
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991