Gaze-directed Visual Grounding under Object Referring Uncertainty
Dailing Zhang, Yinxiao Tian, Kaifeng Chen, Kun Qian
- Year
- 2022
- Citations
- 4
Abstract
Visual Grounding (VG) is a promising approach toward object detection using both visual and text inputs. However, when the text inputs are from natural human language in a human-robot interaction context, the human intents of object referring are usually underspecified. In this case, the robustness and accuracy of existing visual grounding approaches will deteriorate. To this end, in this paper, we propose a Gaze-directed Visual Grounding (GVG) network that incorporates gaze information. The function of uniquely grounding the user's referenced object is achieved by fusing the user's gaze estimation, underspecified reference, and image information in the target scene. To accommodate the errors in gaze estimation, we model the probability distribution of gaze point as a Gaussian distribution, whose parameters are learned from real-world training data. The RefCoco dataset is extended by adding the gaze distribution as masks in images for training the GVG network. We found that the accuracy of the GVG network under underspecified object reference is significantly improved compared to the VG network. Quantitative results show that gaze information increased the practical application of VG in the human-computer interaction process.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002