Papers
1
Total Citations
3
H-Index
1
About
Yulong Ao is a leading researcher in artificial intelligence, with a primary focus on multimodal learning and large-scale generative models. His most influential work centers on advancing next-token prediction frameworks to unify learning across text, images, and video—a fundamental challenge in AI. In his landmark 2026 paper, "Multimodal learning with next-token prediction for large multimodal models," Ao proposed a novel algorithm that extends the success of large language models into multimodal domains, enabling seamless generation and understanding across diverse data types. This contribution has quickly garnered attention, earning 3 citations in its early stages, and is poised to shape the next generation of foundation models. Ao’s research bridges critical gaps between language and vision, offering a scalable, unified approach to multimodal intelligence. His work is particularly notable for its potential to democratize AI capabilities, making sophisticated multimodal interaction more accessible. As a rising figure in the field, Ao’s innovative algorithms are already influencing both academic research and practical applications in generative AI.
Research Focus
Key Achievements
Top Papers
- 1