Home /Research /ANSEL Photobot: A Robot Event Photographer with Semantic Intelligence
OTHER

ANSEL Photobot: A Robot Event Photographer with Semantic Intelligence

Dmitriy Rivkin, Gregory Dudek, Nikhil Kakodkar, David Meger, Oliver Limoyo, Michael Jenkin, Xue Liu, Francois R. Hogan

Year
2023
Citations
7

Abstract

Our work examines the way in which large language models can be used for robotic planning and sampling in the context of automated photographic documentation. Specifically, we illustrate how to produce a photo-taking robot with an exceptional level of semantic awareness by leveraging recent advances in general purpose language (LM) and vision-language (VLM) models. Given a high-level description of an event we use an LM to generate a natural-language list of photo descriptions that one would expect a photographer to capture at the event. We then use a VLM to identify the best matches to these descriptions in the robot's video stream. The photo portfolios generated by our method are consistently rated as more appropriate to the event by human evaluators than those generated by existing methods.

Keywords

Computer scienceEvent (particle physics)RobotDocumentationContext (archaeology)Artificial intelligenceNatural languageHuman–computer interactionSemantics (computer science)Natural language processing

Related papers

Browse all OTHER papers