首页 /研究 /Integrating Vision Foundation Models with Reinforcement Learning for Enhanced Object Interaction

PERCEPTION

Integrating Vision Foundation Models with Reinforcement Learning for Enhanced Object Interaction

Ahmad Farooq, Kamran Iqbal

发表年份: 2025
访问权限: 开放获取

摘要

This paper presents a novel approach that integrates vision foundation models with reinforcement learning to enhance object interaction capabilities in simulated environments. By combining the Segment Anything Model (SAM) and YOLOv5 with a Proximal Policy Optimization (PPO) agent operating in the AI2-THOR simulation environment, we enable the agent to perceive and interact with objects more effectively. Our comprehensive experiments, conducted across four diverse indoor kitchen settings, demonstrate significant improvements in object interaction success rates and navigation efficiency compared to a baseline agent without advanced perception. The results show a 68% increase in average cumulative reward, a 52.5% improvement in object interaction success rate, and a 33% increase in navigation efficiency. These findings highlight the potential of integrating foundation models with reinforcement learning for complex robotic tasks, paving the way for more sophisticated and capable autonomous agents.

关键词

cs.ROcs.AIcs.CVcs.LGeess.SY

Integrating Vision Foundation Models with Reinforcement Learning for Enhanced Object Interaction

摘要

关键词

相关论文

Artificial intelligence: a modern approach

Are we ready for autonomous driving? The KITTI vision benchmark suite

TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems

Vision meets robotics: The KITTI dataset