Visual-Language Decision System Through Integration of Foundation Models for Service Robot Navigation
Peiyuan Zhu, Lotfi El Hafi, Tadahiro Taniguchi
- Year
- 2024
- Citations
- 5
Abstract
This study aims to build a system that bridges the gap between robotics and environmental understanding by integrating various foundation models. While current visual-language models (VLMs) and large language models (LLMs) have demonstrated robust capabilities in image recognition and language comprehension, challenges remain in integrating them into practical robotic applications. Therefore, we propose a visual-language decision (VLD) system that allows a robot to autonomously analyze its surroundings using three VLMs (CLIP, OFA, and PaddleOCR) to generate semantic information. This information is further processed using the GPT-3 LLM, which allows the robot to make judgments during autonomous navigation. The contribution is twofold: 1) We show that integrating CLIP, OFA, and PaddleOCR into a robotic system can generate task-critical information in unexplored environments; 2) We explore how to effectively use GPT-3 to match the results generated by specific VLMs and make navigation decisions based on environmental information. We also implement a photorealistic training environment using Isaac Sim to test and validate the proposed VLD system in simulation. Finally, we demonstrate VLD-based real-world navigation in an unexplored environment using a TurtleBot3 robot equipped with a lidar and an RGB camera.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991