Papers

17

Total Citations

340

H-Index

6

About

Sean Kirmani is a robotics and AI researcher whose work sits at the intersection of vision-language models, spatial reasoning, and robotic learning. His most influential contribution, **SpatialVLM** (2024, 163 citations), addresses a critical gap in modern AI by endowing Vision-Language Models with 3D spatial reasoning capabilities — a foundational requirement for both visual question answering and real-world robotics deployment. Complementing this, his work on **Language to Rewards for Robotic Skill Synthesis** (2023) and **PromptBook for Manipulation Skills** (2024) demonstrates his sustained focus on leveraging large language models to bridge the gap between semantic reasoning and low-level robotic control. Kirmani has also contributed to large-scale robot deployment, including deep reinforcement learning systems for waste sorting across office building fleets and the **AutoRT** framework for orchestrating robotic agents at scale. His earlier work explored human-robot interaction through LED-based robot signaling and semantic mapping for navigation. Across his career, Kirmani has consistently tackled challenges of generalization, scalability, and real-world grounding — problems central to making robots genuinely useful. With over 300 cumulative citations, his research is shaping how the next generation of intelligent, language-guided robots will perceive and act in the physical world.

Research Focus

Key Achievements

6
H-Index
17
Papers
340
Total Citations
20
Avg Citations/Paper
🏆 Most Cited Paper
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
163 citations · 2024
📈 Most Prolific Year: 2024 (6 Papers)
🤝 Key Collaborators: 201
🏛 Institutions: Google (United States), The University of Texas at Austin, Google DeepMind (United Kingdom)

Top Papers

  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7
  8. 8
  9. 9
  10. 10

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 14 days ago