About

Alexey K. Kovalev is a researcher at the forefront of embodied AI and multimodal machine learning, with a focus on bridging the gap between language understanding and physical action. His work centers on developing intelligent agents that can interpret natural language instructions, plan complex behaviors, and adapt to real-world environments. Kovalev’s major contributions include the introduction of the **RozumFormer**, a specialized multimodal transformer for controlling robotic agents in object manipulation tasks, and **LERa** (Look, Explain, Replan), a visual language model-based replanning approach that enables robots to recover from failures using visual feedback. He has also advanced the field through foundational studies on using large language models for embodied planning and common-sense verification. With over 35 citations across his most-cited works, Kovalev’s research has been recognized for its practical impact, particularly in creating the **AmbiK** dataset for ambiguous kitchen tasks. His innovative approaches to task planning, ambiguity detection, and visual-language integration are shaping the next generation of autonomous, instruction-following robots.

Research Focus

Key Achievements

3
H-Index
6
Papers
35
Total Citations
6
Avg Citations/Paper
🏆 Most Cited Paper
Vector Semiotic Model for Visual Question Answering
13 citations · 2021
📈 Most Prolific Year: 2023 (2 Papers)
🤝 Key Collaborators: 21
🏛 Institutions: National Research University Higher School of Economics, Moscow Institute of Physics and Technology, United States Air Force

Top Papers

  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 14 days ago