Jiaming Song
Papers
2
Total Citations
9
H-Index
2
About
Jiaming Song is a leading researcher at the intersection of computer vision, natural language processing, and robotics, with a primary focus on enabling machines to understand and predict complex, language-guided behaviors. His work centers on developing generative models that bridge the gap between high-level human instructions and low-level robotic control. Song’s major contributions include pioneering text-conditioned video prediction (TVP) with his paper "Seer: Language Instructed Video Prediction with Latent Diffusion Models" (2023, 6 citations), which empowers robots to foresee future trajectories from language commands—a critical capability for robust planning and goal achievement. He also introduced "LISA: Learning Interpretable Skill Abstractions from Language" (2022, 3 citations), a framework for learning reusable, interpretable skills from language instructions, enabling more effective generalization in multi-task environments. By tackling the challenge of grounding language in visual and motor predictions, Song’s work is shaping the next generation of generalist robots that can understand and act upon human commands in dynamic, real-world settings.
Research Focus
Key Achievements
Top Papers
- 1Seer: Language Instructed Video Prediction with Latent Diffusion Models6 citations · 2023
- 2LISA: Learning Interpretable Skill Abstractions from Language3 citations · 2022