Moo Jin Kim
Papers
5
Total Citations
123
H-Index
5
About
Moo Jin Kim is a rising star in robot learning, whose work bridges computer vision and manipulation to make robots more capable and data-efficient. His research centers on vision-language-action models (VLAs), novel-view synthesis for robotics, and scalable dataset creation. Kim made a foundational contribution with **OpenVLA** (2024, 39 citations), an open-source model that combines Internet-scale vision-language pretraining with diverse robot demonstration data, enabling robust policy fine-tuning without training from scratch. He further advanced this paradigm with **CoT-VLA** (2025, 23 citations), introducing visual chain-of-thought reasoning to improve VLA generalization. In **NeRF in the Palm of Your Hand** (2023, 39 citations), Kim pioneered corrective augmentation via novel-view synthesis, reducing the need for large expert demonstrations in imitation learning. He also co-created **BridgeData V2** (2023, 12 citations), a large-scale, diverse manipulation dataset with over 60,000 trajectories that has become a key benchmark for scalable robot learning. His earlier work on hand-centric visual perspectives (2022, 10 citations) challenged conventional camera placements, showing that eye-in-hand views can improve generalization despite reduced observability. With multiple highly-cited papers in just a few years, Kim is shaping how robots learn from vision and language.
Research Focus
Key Achievements
Top Papers
- 1
- 2OpenVLA: An Open-Source Vision-Language-Action Model39 citations · 2024
- 3CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models23 citations · 2025
- 4BridgeData V2: A Dataset for Robot Learning at Scale12 citations · 2023
- 5Vision-Based Manipulators Need to Also See from Their Hands10 citations · 2022