首页 /研究 /Interpretable and Globally Optimal Prediction for Textual Grounding\n using Image Concepts
OTHER

Interpretable and Globally Optimal Prediction for Textual Grounding\n using Image Concepts

Raymond A. Yeh, Jinjun Xiong, Wen‐mei Hwu, Nguyen Q. Minh, Alexander G. Schwing

发表年份
2018
引用次数
42
访问权限
开放获取

摘要

Textual grounding is an important but challenging task for human-computer\ninteraction, robotics and knowledge mining. Existing algorithms generally\nformulate the task as selection from a set of bounding box proposals obtained\nfrom deep net based systems. In this work, we demonstrate that we can cast the\nproblem of textual grounding into a unified framework that permits efficient\nsearch over all possible bounding boxes. Hence, the method is able to consider\nsignificantly more proposals and doesn't rely on a successful first stage\nhypothesizing bounding box proposals. Beyond, we demonstrate that the trained\nparameters of our model can be used as word-embeddings which capture\nspatial-image relationships and provide interpretability. Lastly, at the time\nof submission, our approach outperformed the current state-of-the-art methods\non the Flickr 30k Entities and the ReferItGame dataset by 3.08% and 7.77%\nrespectively.\n

关键词

InterpretabilityBounding overwatchComputer scienceMinimum bounding boxTask (project management)Artificial intelligenceSelection (genetic algorithm)Set (abstract data type)Word (group theory)Machine learning

相关论文

查看 OTHER 分类全部论文