Gregory Shakhnarovich
Papers
6
Total Citations
93
H-Index
3
About
Gregory Shakhnarovich is a leading researcher at the intersection of computer vision, speech processing, and natural language understanding, with a particular focus on visually grounded learning. His pioneering work has fundamentally advanced how machines can learn from unlabelled speech paired with visual context—a paradigm critical for low-resource speech processing, robotics, and modeling human language acquisition. Shakhnarovich’s most influential contributions include developing models that map images and spoken captions into a shared semantic space, enabling tasks like semantic speech retrieval and keyword spotting without transcribed data. His 2017 paper on "Semantic speech retrieval with a visually grounded model of untranscribed speech" (50 citations) remains a cornerstone in this area. Beyond speech, he has made notable contributions to neural decoding for motor prostheses, including work on decoding grasp aperture from motor-cortical activity (2007, 24 citations). Most recently, his 2024 work on "Transcrib3D" tackles the challenge of interpreting 3D referring expressions using large language models, a critical step for human-robot interaction. Shakhnarovich’s research consistently bridges perception and language, driving progress toward more intuitive, multimodal AI systems.
Research Focus
Key Achievements
Top Papers
- 1
- 2Decoding grasp aperture from motor-cortical population activity24 citations · 2007
- 3Visually Grounded Learning of Keyword Prediction from Untranscribed Speech12 citations · 2017
- 4
- 5Semantic keyword spotting by learning from images and speech.2 citations · 2017
- 6