Herman Kamper
Papers
3
Total Citations
64
H-Index
2
About
Herman Kamper is a leading researcher in visually grounded speech processing, a field at the intersection of machine learning, natural language processing, and cognitive science. His work focuses on developing models that learn directly from untranscribed speech paired with visual context—an approach inspired by human language acquisition. Kamper’s major contributions include pioneering methods for semantic speech retrieval and keyword prediction without the need for textual transcriptions. His 2017 paper, "Semantic speech retrieval with a visually grounded model of untranscribed speech," has garnered 50 citations, establishing a foundational framework for learning from unlabelled audio-visual data. This work is particularly impactful for low-resource speech processing, robotics, and understanding how infants acquire language. Kamper further advanced the field with "Visually Grounded Learning of Keyword Prediction from Untranscribed Speech" (12 citations), demonstrating how robots and machines can leverage visual cues to ground spoken language. His research has significant implications for building more robust, human-like AI systems that can learn in resource-constrained environments. Kamper’s innovative approach continues to shape the future of speech technology and multimodal learning.
Research Focus
Key Achievements
Top Papers
- 1
- 2Visually Grounded Learning of Keyword Prediction from Untranscribed Speech12 citations · 2017
- 3Semantic keyword spotting by learning from images and speech.2 citations · 2017