Jiasen Lu
Papers
3
Total Citations
60
H-Index
2
About
Jiasen Lu is a leading researcher in multimodal AI, whose work centers on building unified models that seamlessly integrate vision, language, audio, and action. His most impactful contribution is **Unified-IO 2**, the first autoregressive multimodal model capable of both understanding and generating across these four modalities. By tokenizing diverse inputs—from images and text to bounding boxes and action sequences—into a shared semantic space, Lu’s architecture enables a single model to perform tasks ranging from image captioning to robotic control. This breakthrough, detailed in his 2024 paper (55 citations), represents a major step toward general-purpose AI agents. Earlier, Lu explored **visual curiosity** (2018), developing agents that proactively ask questions to learn about unrecognized objects—a foundational idea for interactive, lifelong learning systems. His work bridges the gap between perception and reasoning, with implications for robotics, accessibility, and human-computer interaction. With a growing citation footprint and a focus on scaling multimodal models, Lu is shaping the future of AI that can perceive, communicate, and act in the world.
Research Focus
Key Achievements
Top Papers
- 1
- 2
- 3Visual Curiosity: Learning to Ask Questions to Learn Visual Recognition2 citations · 2018