Jiaming Song

Papers

2

Total Citations

9

H-Index

2

About

Jiaming Song is a leading researcher at the intersection of computer vision, natural language processing, and robotics, with a primary focus on enabling machines to understand and predict complex, language-guided behaviors. His work centers on developing generative models that bridge the gap between high-level human instructions and low-level robotic control. Song’s major contributions include pioneering text-conditioned video prediction (TVP) with his paper "Seer: Language Instructed Video Prediction with Latent Diffusion Models" (2023, 6 citations), which empowers robots to foresee future trajectories from language commands—a critical capability for robust planning and goal achievement. He also introduced "LISA: Learning Interpretable Skill Abstractions from Language" (2022, 3 citations), a framework for learning reusable, interpretable skills from language instructions, enabling more effective generalization in multi-task environments. By tackling the challenge of grounding language in visual and motor predictions, Song’s work is shaping the next generation of generalist robots that can understand and act upon human commands in dynamic, real-world settings.

Research Focus

Key Achievements

2
H-Index
2
Papers
9
Total Citations
5
Avg Citations/Paper
🏆 Most Cited Paper
Seer: Language Instructed Video Prediction with Latent Diffusion Models
6 citations · 2023
📈 Most Prolific Year: 2023 (1 Papers)
🤝 Key Collaborators: 6

Top Papers

  1. 1
  2. 2

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 14 days ago