Xuebin Min

Beijing Academy of Artificial Intelligence

Papers

1

Total Citations

3

H-Index

1

About

Xuebin Min is a leading researcher in artificial intelligence, with a primary focus on multimodal learning and large-scale generative models. His most influential work tackles the fundamental challenge of developing unified algorithms capable of learning from and generating across diverse modalities—including text, images, and video. Min is best known for pioneering the extension of next-token prediction—a cornerstone of large language models—into the multimodal domain. His 2026 paper on this topic, which has already garnered 3 citations, proposes a novel framework that enables a single model to seamlessly process and produce content across visual and textual formats. This contribution addresses a critical bottleneck in AI, moving beyond modality-specific architectures toward truly integrated understanding and generation. By demonstrating that the same predictive principle driving language models can be effectively applied to vision and video, Min’s work has opened new pathways for building more versatile and efficient multimodal systems. His research holds significant promise for applications in autonomous content creation, human-computer interaction, and accessible AI, establishing him as an emerging voice in the next wave of foundational model design.

Research Focus

Key Achievements

1
H-Index
1
Papers
3
Total Citations
3
Avg Citations/Paper
🏆 Most Cited Paper
Multimodal learning with next-token prediction for large multimodal models
3 citations · 2026
📈 Most Prolific Year: 2026 (1 Papers)
🤝 Key Collaborators: 24
🏛 Institutions: Beijing Academy of Artificial Intelligence

Top Papers

  1. 1

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 11 days ago