首页 /研究 /Semi-supervised vision-language mapping via variational learning
OTHER

Semi-supervised vision-language mapping via variational learning

Yuming Shen, Li Zhang, Ling Shao

发表年份
2017
引用次数
5

摘要

Understanding the semantic relations between vision and language data has become a research trend in artificial intelligence and robotic systems. The lack of training data is an essential issue for vision-language understanding. We address the problem of image and sentence cross-modal retrieval when paired training samples are not sufficient. Inspired by recent works in variational inference, in this paper, the autoencoding variational Bayes framework is novelly extended to a semi-supervised model for image-sentence mapping task. Our method does not require all training images and sentences to be paired. The proposed model is an end-to-end system, and consists of a two-level variational embedding structure where unpaired data are involved in the first level embedding to give support to intra-modality statistics so that the lower bound of the joint marginal likelihood of paired data embeddings can be better approximated. The proposed retrieval model is evaluated on two popular datasets, i.e. Flickr30K and Flickr8K, producing superior performances compared with related state-of-the-art methods.

关键词

Computer scienceEmbeddingArtificial intelligenceSentenceInferenceImage (mathematics)Natural language processingMachine learningPattern recognition (psychology)

相关论文

查看 OTHER 分类全部论文