Suman Saha
Papers
6
Total Citations
52
H-Index
5
About
Suman Saha is a computer vision researcher whose work centers on human action detection, video understanding, and visual grounding. His most significant contributions lie in developing systems capable of detecting and localizing human actions in video through the construction of "action tubes" — spatiotemporal sequences that track actions across frames. His 2019 paper "Predicting Action Tubes," his most cited work with 19 citations, advanced the field by enabling anticipatory action detection, while his series of papers on incremental tube construction addressed a critical gap in real-time, online action detection systems suited for applications like human-robot interaction. His 2020 work on Two-Stream AMTnet further pushed the state of the art by integrating video-based action representations with incremental tube generation, moving beyond the limitations of frame-level detection. Saha also contributed to autonomous driving research through the READ dataset, bridging action detection with real-world vehicular perception challenges. More recently, his work has expanded into 3D visual grounding and verbo-visual fusion, reflecting a broadening research vision. With publications spanning nearly a decade, Saha has established himself as a consistent contributor to the intersection of video analysis, human activity recognition, and embodied AI applications.
Research Focus
Key Achievements
Top Papers
- 1Predicting Action Tubes19 citations · 2019
- 2Incremental Tube Construction for Human Action Detection12 citations · 2017
- 3Two-Stream AMTnet for Action Detection8 citations · 2020
- 4Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding5 citations · 2024
- 5Incremental Tube Construction for Human Action Detection5 citations · 2018
- 6Action Detection from a Robot-Car Perspective3 citations · 2018