Papers
20
Total Citations
406
H-Index
11
About
Osbert Bastani is a prominent researcher at the intersection of reinforcement learning, robot learning, and AI safety, whose work has significantly advanced how intelligent systems learn complex behaviors while operating reliably in the real world. His research spans safe reinforcement learning, vision-language reward learning, and large-scale robot manipulation, addressing some of the field's most pressing challenges. Bastani's contributions to safe RL are particularly noteworthy. His model predictive shielding frameworks — applied to both single-agent and multi-agent settings — offer principled guarantees against unsafe behaviors during policy learning, accumulating nearly 50 citations combined. He has also pioneered reward and representation learning for robotics, with projects like VIP and LIV demonstrating how human videos and language supervision can unlock scalable robot skill acquisition without costly task-specific data. More recently, Bastani contributed to Eureka, a compelling demonstration that large language models can autonomously design reward functions for dexterous manipulation tasks at human-level quality (48 citations), and to DROID, a landmark large-scale robot manipulation dataset already garnering 108 citations since 2024. Beyond robotics, his work extends to clinical AI, including applications in glaucoma diagnosis. Collectively, his research reflects a rare breadth — pushing frontiers in both foundational methodology and real-world deployment.
Research Focus
Key Achievements
Top Papers
- 1DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset108 citations · 2024
- 2Eureka: Human-Level Reward Design via Coding Large Language Models48 citations · 2023
- 3
- 4
- 5Safe Reinforcement Learning via Statistical Model Predictive Shielding25 citations · 2021
- 6
- 7LIV: Language-Image Representations and Rewards for Robotic Control24 citations · 2023
- 8A Composable Specification Language for Reinforcement Learning Tasks21 citations · 2020
- 9Conservative Offline Distributional Reinforcement Learning15 citations · 2021
- 10Learning Safe Unlabeled Multi-Robot Planning with Motion Constraints13 citations · 2019