Makarand Tapaswi
University of Toronto, Indian Institute of Technology Hyderabad
Papers
4
Total Citations
27
H-Index
3
About
Makarand Tapaswi is a researcher working at the intersection of computer vision, robotics, and multimodal learning, with a particular focus on enabling machines to understand and interact with the physical and social world. His most recognized contribution, "Learning Object Manipulation Skills via Real Videos" (2020, 13 citations), pioneered an approach to robot learning that bypasses costly expert demonstrations by leveraging instructional video footage — a significant step toward more accessible robot training pipelines. Building on this, his 2022 work on instruction-driven, history-aware policies for robotic manipulation (7 citations) advanced the field by integrating natural language understanding with long-horizon motor control, addressing critical challenges in task generalization. Beyond robotics, Tapaswi has made meaningful contributions to social scene understanding through MovieGraphs (2018, 4 citations), a dataset designed to help AI systems decode human emotions, motivations, and interpersonal dynamics from film. His more recent work on audio-visual physics inference from pouring liquids (2025, 3 citations) reflects a broader intellectual curiosity, connecting sensory perception to physical reasoning. Together, his research paints a picture of a scientist devoted to grounding artificial intelligence in the rich complexity of human experience.
Research Focus
Key Achievements
Top Papers
- 1
- 2Instruction-driven history-aware policies for robotic manipulations7 citations · 2022
- 3MovieGraphs: Towards Understanding Human-Centric Situations from Videos4 citations · 2018
- 4The Sound of Water: Inferring Physical Properties from Pouring Liquids3 citations · 2025