Interpretable and Globally Optimal Prediction for Textual Grounding\n using Image Concepts
Raymond A. Yeh, Jinjun Xiong, Wen‐mei Hwu, Nguyen Q. Minh, Alexander G. Schwing
- Year
- 2018
- Citations
- 42
- Access
- Open access
Abstract
Textual grounding is an important but challenging task for human-computer\ninteraction, robotics and knowledge mining. Existing algorithms generally\nformulate the task as selection from a set of bounding box proposals obtained\nfrom deep net based systems. In this work, we demonstrate that we can cast the\nproblem of textual grounding into a unified framework that permits efficient\nsearch over all possible bounding boxes. Hence, the method is able to consider\nsignificantly more proposals and doesn't rely on a successful first stage\nhypothesizing bounding box proposals. Beyond, we demonstrate that the trained\nparameters of our model can be used as word-embeddings which capture\nspatial-image relationships and provide interpretability. Lastly, at the time\nof submission, our approach outperformed the current state-of-the-art methods\non the Flickr 30k Entities and the ReferItGame dataset by 3.08% and 7.77%\nrespectively.\n
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991