首页 /研究 /Interactive Visual Grounding of Referring Expressions for Human-Robot\n Interaction
HRI

Interactive Visual Grounding of Referring Expressions for Human-Robot\n Interaction

M.H Shridhar, David Hsu

发表年份
2018
引用次数
6
访问权限
开放获取

摘要

This paper presents INGRESS, a robot system that follows human natural\nlanguage instructions to pick and place everyday objects. The core issue here\nis the grounding of referring expressions: infer objects and their\nrelationships from input images and language expressions. INGRESS allows for\nunconstrained object categories and unconstrained language expressions.\nFurther, it asks questions to disambiguate referring expressions interactively.\nTo achieve these, we take the approach of grounding by generation and propose a\ntwo-stage neural network model for grounding. The first stage uses a neural\nnetwork to generate visual descriptions of objects, compares them with the\ninput language expression, and identifies a set of candidate objects. The\nsecond stage uses another neural network to examine all pairwise relations\nbetween the candidates and infers the most likely referred object. The same\nneural networks are used for both grounding and question generation for\ndisambiguation. Experiments show that INGRESS outperformed a state-of-the-art\nmethod on the RefCOCO dataset and in robot experiments with humans.\n

关键词

Computer sciencePairwise comparisonArtificial intelligenceGroundObject (grammar)RobotSet (abstract data type)Artificial neural networkNatural languageExpression (computer science)

相关论文

查看 HRI 分类全部论文