Savya Khosla

University of Illinois Urbana-Champaign

Papers

2

Total Citations

58

H-Index

2

About

Savya Khosla is a leading researcher in multimodal AI, whose work centers on unifying vision, language, audio, and action within a single autoregressive framework. Their major contribution is the development of Unified-IO 2, the first model capable of both understanding and generating across these four modalities—a breakthrough that eliminates the need for separate, task-specific systems. By tokenizing diverse inputs like images, text, audio, and bounding boxes into a shared semantic space, Khosla’s approach enables a single model to handle everything from image captioning to action prediction. This work has already garnered over 55 citations in its first year, signaling its rapid impact on the field. Khosla’s achievement lies in scaling multimodal learning without sacrificing generality, paving the way for more versatile AI assistants. Their research is foundational for students and engineers aiming to build systems that perceive and interact with the world as humans do—through multiple senses and actions simultaneously.

Research Focus

Key Achievements

2
H-Index
2
Papers
58
Total Citations
29
Avg Citations/Paper
🏆 Most Cited Paper
Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
55 citations · 2024
📈 Most Prolific Year: 2024 (1 Papers)
🤝 Key Collaborators: 8
🏛 Institutions: University of Illinois Urbana-Champaign

Top Papers

  1. 1
  2. 2

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 12 days ago