Minghuan Liu
Papers
6
Total Citations
44
H-Index
3
About
Minghuan Liu is an emerging researcher at the intersection of robotics, computer vision, and foundation models, with a focus on enabling more capable, generalizable robotic systems through large-scale pre-training and multimodal learning. His work explores how vision-language models (VLMs) and generative pre-trained architectures can be effectively adapted for robotic manipulation and control, as demonstrated in his influential 2023 paper "Vision-Language Foundation Models as Effective Robot Imitators" (19 citations), which established a streamlined framework for fine-tuning VLMs for imitation learning in robotics. Liu has also made notable contributions to depth estimation, introducing the innovative "Prompt Depth Anything" paradigm for 4K-resolution metric depth estimation (15 citations), extending the prompting methodology from language and vision into depth foundation models. His broader research agenda encompasses legged robot loco-manipulation with whole-body control, scalable simulation data generation via large language models (GenSim2), and the principled design of vision-language-action models for generalist robot policies. Together, his portfolio reflects a coherent mission: bridging powerful foundation models with real-world robotic intelligence, making him a researcher to watch in embodied AI.
Research Focus
Key Achievements
Top Papers
- 1Vision-Language Foundation Models as Effective Robot Imitators19 citations · 2023
- 2Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation15 citations · 2025
- 3
- 4Visual Whole-Body Control for Legged Loco-Manipulation2 citations · 2024
- 5
- 6