Papers

1

Total Citations

2

H-Index

1

About

Zhen Xing is a leading researcher in computer vision and generative AI, with a primary focus on video prediction and multimodal diffusion models. Their most notable contribution is the development of "Aid," a pioneering framework that adapts Image2Video diffusion models for instruction-guided video prediction. This work addresses the critical challenge of Text-guided Video Prediction (TVP), where future frames are generated from an initial frame based on natural language instructions—a capability with transformative applications in virtual reality, robotics, and automated content creation. By successfully adapting Stable Diffusion for this task, Xing has advanced the state of the art in controllable video generation, enabling more intuitive and precise motion synthesis. Their research bridges the gap between static image understanding and dynamic video generation, offering a practical solution for real-world scenarios requiring temporal reasoning. With early citations already accumulating, Xing's work is poised to influence both academic research and industrial applications, establishing them as an emerging authority in the intersection of diffusion models and video prediction.

Research Focus

Key Achievements

1
H-Index
1
Papers
2
Total Citations
2
Avg Citations/Paper
🏆 Most Cited Paper
Aid: Adapting Image2video Diffusion Models for Instruction-Guided Video Prediction
2 citations · 2025
📈 Most Prolific Year: 2025 (1 Papers)
🤝 Key Collaborators: 4
🏛 Institutions: Shanghai Key Laboratory of Trustworthy Computing

Top Papers

  1. 1

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 10 days ago