首页 /研究 /Lan-grasp: Using Large Language Models for Semantic Object Grasping and Placement
MANIPULATION

Lan-grasp: Using Large Language Models for Semantic Object Grasping and Placement

Reihaneh Mirjalili, Michael Krawez, Yannik Blei, Simone Silenzi, Wolfram Burgard

发表年份
2023
引用次数
5
访问权限
开放获取

摘要

In this paper, we propose Lan-grasp, a novel approach towards more appropriate semantic grasping and placing. We leverage foundation models to equip the robot with a semantic understanding of object geometry, enabling it to identify the right place to grasp, which parts to avoid, and the natural pose for placement. This is an important contribution to grasping and utilizing objects in a more meaningful and safe manner. We leverage a combination of a Large Language Model, a Vision-Language Model, and a traditional grasp planner to generate grasps that demonstrate a deeper semantic understanding of the objects. Building on foundation models provides us with a zero-shot grasp method that can handle a wide range of objects without requiring further training or fine-tuning. We also propose a method for safely putting down a grasped object. The core idea is to rotate the object upright utilizing a pretrained generative model and the reasoning capabilities of a VLM. We evaluate our method in real-world experiments on a custom object dataset and present the results of a survey that asks participants to choose an object part appropriate for grasping. The results show that the grasps generated by our method are consistently ranked higher by the participants than those generated by a conventional grasping planner and a recent semantic grasping approach. In addition, we propose a Visual Chain-of-Thought feedback loop to assess grasp feasibility in complex scenarios. This mechanism enables dynamic reasoning and generates alternative grasp strategies when needed, ensuring safer and more effective grasping outcomes.

关键词

GRASPComputer scienceLeverage (statistics)Object (grammar)Artificial intelligencePlannerSet (abstract data type)Computer visionRobotHuman–computer interaction

相关论文

查看 MANIPULATION 分类全部论文