Suman Saha

Oxford Brookes University, ETH Zurich

Papers

6

Total Citations

52

H-Index

5

About

Suman Saha is a computer vision researcher whose work centers on human action detection, video understanding, and visual grounding. His most significant contributions lie in developing systems capable of detecting and localizing human actions in video through the construction of "action tubes" — spatiotemporal sequences that track actions across frames. His 2019 paper "Predicting Action Tubes," his most cited work with 19 citations, advanced the field by enabling anticipatory action detection, while his series of papers on incremental tube construction addressed a critical gap in real-time, online action detection systems suited for applications like human-robot interaction. His 2020 work on Two-Stream AMTnet further pushed the state of the art by integrating video-based action representations with incremental tube generation, moving beyond the limitations of frame-level detection. Saha also contributed to autonomous driving research through the READ dataset, bridging action detection with real-world vehicular perception challenges. More recently, his work has expanded into 3D visual grounding and verbo-visual fusion, reflecting a broadening research vision. With publications spanning nearly a decade, Saha has established himself as a consistent contributor to the intersection of video analysis, human activity recognition, and embodied AI applications.

Research Focus

Key Achievements

5
H-Index
6
Papers
52
Total Citations
9
Avg Citations/Paper
🏆 Most Cited Paper
Predicting Action Tubes
19 citations · 2019
📈 Most Prolific Year: 2018 (2 Papers)
🤝 Key Collaborators: 11
🏛 Institutions: Oxford Brookes University, ETH Zurich

Top Papers

  1. 1
    Predicting Action Tubes
    19 citations · 2019
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 14 days ago