Norman Di Palo
Papers
10
Total Citations
142
H-Index
6
About
Norman Di Palo is at the forefront of a new wave in robotics, pioneering the use of foundation models—large language models (LLMs) and vision transformers—to create more intelligent, sample-efficient, and generalizable robot learning systems. His research centers on imitation learning, in-context learning, and safe human-robot interaction, with a core focus on enabling robots to learn complex manipulation tasks from minimal human demonstrations. Di Palo’s major contributions include demonstrating that LLMs can act as zero-shot trajectory generators (46 citations), challenging the assumption that they are only suitable for high-level planning. He also introduced DINOBot (26 citations), a novel imitation learning framework that leverages DINOv2 features for robust visual retrieval and alignment, and developed a method for in-context imitation learning using off-the-shelf text-based transformers (20 citations). His work on learning multi-stage tasks from a single demonstration and the SAFARI algorithm for safe active imitation learning further highlights his impact. With over 160 total citations and recent contributions to the Gemini Robotics project, Di Palo is a rising leader in building unified, foundation-model-driven robotic agents that bring AI into the physical world.
Research Focus
Key Achievements
Top Papers
- 1Language Models as Zero-Shot Trajectory Generators46 citations · 2024
- 2
- 3Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics20 citations · 2024
- 4Towards A Unified Agent with Foundation Models17 citations · 2023
- 5On the Effectiveness of Retrieval, Alignment, and Replay in Manipulation10 citations · 2024
- 6Learning Multi-Stage Tasks with One Demonstration via Self-Replay8 citations · 2021
- 7SAFARI: Safe and Active Robot Imitation Learning with Imagination4 citations · 2020
- 8Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics4 citations · 2024
- 9Gemini Robotics: Bringing AI into the Physical World4 citations · 2025
- 10R+X: Retrieval and Execution from Everyday Human Videos3 citations · 2025