Minghuan Liu

Shanghai Jiao Tong University

Papers

6

Total Citations

44

H-Index

3

About

Minghuan Liu is an emerging researcher at the intersection of robotics, computer vision, and foundation models, with a focus on enabling more capable, generalizable robotic systems through large-scale pre-training and multimodal learning. His work explores how vision-language models (VLMs) and generative pre-trained architectures can be effectively adapted for robotic manipulation and control, as demonstrated in his influential 2023 paper "Vision-Language Foundation Models as Effective Robot Imitators" (19 citations), which established a streamlined framework for fine-tuning VLMs for imitation learning in robotics. Liu has also made notable contributions to depth estimation, introducing the innovative "Prompt Depth Anything" paradigm for 4K-resolution metric depth estimation (15 citations), extending the prompting methodology from language and vision into depth foundation models. His broader research agenda encompasses legged robot loco-manipulation with whole-body control, scalable simulation data generation via large language models (GenSim2), and the principled design of vision-language-action models for generalist robot policies. Together, his portfolio reflects a coherent mission: bridging powerful foundation models with real-world robotic intelligence, making him a researcher to watch in embodied AI.

Research Focus

Key Achievements

3
H-Index
6
Papers
44
Total Citations
7
Avg Citations/Paper
🏆 Most Cited Paper
Vision-Language Foundation Models as Effective Robot Imitators
19 citations · 2023
📈 Most Prolific Year: 2023 (2 Papers)
🤝 Key Collaborators: 38
🏛 Institutions: Shanghai Jiao Tong University

Top Papers

  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 14 days ago