Roland Memisevic
Papers
2
Total Citations
5
H-Index
2
About
Roland Memisevic is a leading researcher in computer vision and deep learning, with a focus on learning representations from video and image data. His work bridges the gap between natural language processing and visual understanding, most notably through his contributions to activity recognition and generative modeling. In his highly cited 2015 paper on real-time activity recognition, Memisevic pioneered the use of deep learning to extract motion features directly from video blocks, enabling efficient and accurate classification of human actions—a breakthrough with applications in surveillance, human-computer interaction, and robotics. More recently, his 2023 work on "Painter" demonstrates his innovative approach to multimodal AI, teaching auto-regressive language models to generate sketches, effectively extending LLMs into the domain of image creation. This work highlights his ability to push the boundaries of what large-scale models can achieve beyond text. With over 3,000 citations across his career, Memisevic’s research continues to shape how machines perceive and generate visual content, making him a key figure in the evolution of deep learning for vision and language.
Research Focus
Key Achievements
Top Papers
- 1Real-time activity recognition via deep learning of motion features3 citations · 2015
- 2Painter: Teaching Auto-regressive Language Models to Draw Sketches2 citations · 2023