Jiasen Lu

Allen Institute

Papers

3

Total Citations

60

H-Index

2

About

Jiasen Lu is a leading researcher in multimodal AI, whose work centers on building unified models that seamlessly integrate vision, language, audio, and action. His most impactful contribution is **Unified-IO 2**, the first autoregressive multimodal model capable of both understanding and generating across these four modalities. By tokenizing diverse inputs—from images and text to bounding boxes and action sequences—into a shared semantic space, Lu’s architecture enables a single model to perform tasks ranging from image captioning to robotic control. This breakthrough, detailed in his 2024 paper (55 citations), represents a major step toward general-purpose AI agents. Earlier, Lu explored **visual curiosity** (2018), developing agents that proactively ask questions to learn about unrecognized objects—a foundational idea for interactive, lifelong learning systems. His work bridges the gap between perception and reasoning, with implications for robotics, accessibility, and human-computer interaction. With a growing citation footprint and a focus on scaling multimodal models, Lu is shaping the future of AI that can perceive, communicate, and act in the world.

Research Focus

Key Achievements

2
H-Index
3
Papers
60
Total Citations
20
Avg Citations/Paper
🏆 Most Cited Paper
Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
55 citations · 2024
📈 Most Prolific Year: 2024 (1 Papers)
🤝 Key Collaborators: 12
🏛 Institutions: Allen Institute

Top Papers

  1. 1
  2. 2
  3. 3

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 12 days ago