Yueze Wang

Beijing Academy of Artificial Intelligence

Papers

1

Total Citations

3

H-Index

1

About

Yueze Wang is a leading researcher in multimodal artificial intelligence, with a primary focus on developing unified learning frameworks that bridge text, image, and video modalities. His most cited work, "Multimodal learning with next-token prediction for large multimodal models" (2026), addresses one of the field’s most fundamental challenges: extending the transformative success of next-token prediction—the engine behind large language models—to multimodal domains. This contribution proposes a novel algorithm capable of both learning from and generating across diverse data types, marking a significant step toward truly integrated AI systems. While his citation count is still growing, reflecting the recency of his work, the paper’s conceptual importance has already garnered attention. Wang’s research stands out for tackling the core tension between modality-specific processing and the scalability of autoregressive models, offering a pathway to more versatile and efficient multimodal architectures. His work is particularly relevant for students and researchers interested in the next generation of foundation models, where seamless cross-modal understanding and generation remain open frontiers.

Research Focus

Key Achievements

1
H-Index
1
Papers
3
Total Citations
3
Avg Citations/Paper
🏆 Most Cited Paper
Multimodal learning with next-token prediction for large multimodal models
3 citations · 2026
📈 Most Prolific Year: 2026 (1 Papers)
🤝 Key Collaborators: 24
🏛 Institutions: Beijing Academy of Artificial Intelligence

Top Papers

  1. 1

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 12 days ago