Interpretable and Globally Optimal Prediction for Textual Grounding\n using Image Concepts
Raymond A. Yeh, Jinjun Xiong, Wen‐mei Hwu, Nguyen Q. Minh, Alexander G. Schwing
- 发表年份
- 2018
- 引用次数
- 42
- 访问权限
- 开放获取
摘要
Textual grounding is an important but challenging task for human-computer\ninteraction, robotics and knowledge mining. Existing algorithms generally\nformulate the task as selection from a set of bounding box proposals obtained\nfrom deep net based systems. In this work, we demonstrate that we can cast the\nproblem of textual grounding into a unified framework that permits efficient\nsearch over all possible bounding boxes. Hence, the method is able to consider\nsignificantly more proposals and doesn't rely on a successful first stage\nhypothesizing bounding box proposals. Beyond, we demonstrate that the trained\nparameters of our model can be used as word-embeddings which capture\nspatial-image relationships and provide interpretability. Lastly, at the time\nof submission, our approach outperformed the current state-of-the-art methods\non the Flickr 30k Entities and the ReferItGame dataset by 3.08% and 7.77%\nrespectively.\n
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991