Papers

1

Total Citations

3

H-Index

1

About

Boya Wu is a leading researcher in artificial intelligence, specializing in multimodal learning, large language models, and generative AI. Their most notable contribution is pioneering the extension of next-token prediction—the core mechanism behind large language models—into multimodal domains, enabling unified algorithms that can learn from and generate across text, images, and video. This work, detailed in their highly cited 2026 paper "Multimodal learning with next-token prediction for large multimodal models" (3 citations), addresses a fundamental challenge in AI: creating systems that seamlessly process and produce diverse data types. Wu’s research bridges the gap between language and vision, advancing the development of more versatile and capable multimodal models. Their innovative approach has significant implications for applications ranging from automated content creation to human-computer interaction. With a focus on scalable, unified architectures, Boya Wu continues to shape the future of AI, making their work essential reading for students and researchers exploring the frontiers of multimodal intelligence.

Research Focus

Key Achievements

1
H-Index
1
Papers
3
Total Citations
3
Avg Citations/Paper
🏆 Most Cited Paper
Multimodal learning with next-token prediction for large multimodal models
3 citations · 2026
📈 Most Prolific Year: 2026 (1 Papers)
🤝 Key Collaborators: 24
🏛 Institutions: Beijing Academy of Artificial Intelligence

Top Papers

  1. 1

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 12 days ago