Home /Research /Embodied BERT: A Transformer Model for Embodied, Language-guided Visual\n Task Completion
OTHER

Embodied BERT: A Transformer Model for Embodied, Language-guided Visual\n Task Completion

Alessandro Suglia, Qiaozi Gao, Jesse Thomason, Govind Thattai, Gaurav S. Sukhatme

Year
2021
Citations
30
Access
Open access

Abstract

Language-guided robots performing home and office tasks must navigate in and\ninteract with the world. Grounding language instructions against visual\nobservations and actions to take in an environment is an open challenge. We\npresent Embodied BERT (EmBERT), a transformer-based model which can attend to\nhigh-dimensional, multi-modal inputs across long temporal horizons for\nlanguage-conditioned task completion. Additionally, we bridge the gap between\nsuccessful object-centric navigation models used for non-interactive agents and\nthe language-guided visual task completion benchmark, ALFRED, by introducing\nobject navigation targets for EmBERT training. We achieve competitive\nperformance on the ALFRED benchmark, and EmBERT marks the first\ntransformer-based model to successfully handle the long-horizon, dense,\nmulti-modal histories of ALFRED, and the first ALFRED model to utilize\nobject-centric navigation targets.\n

Keywords

Embodied cognitionTransformerComputer scienceRobotTask (project management)Language understandingBenchmark (surveying)Artificial intelligenceLanguage modelBridge (graph theory)

Related papers

Browse all OTHER papers