Artem Zholus
Papers
1
Total Citations
3
H-Index
1
About
Artem Zholus is an AI researcher working at the frontier of self-supervised learning, world modeling, and embodied intelligence. His work centers on a fundamental question in artificial intelligence: how can machines learn to understand and interact with the physical world primarily through observation, rather than relying on exhaustive labeled data or dense reward signals? His most notable contribution to date is his involvement in **V-JEPA 2** (2025), a landmark self-supervised video modeling framework developed to tackle the challenge of building AI systems capable of understanding, prediction, and physical planning. The approach is notable for its ambitious scope — leveraging internet-scale video data combined with limited robot interaction trajectories to train models that generalize across perception and action. Though early in its citation trajectory with 3 citations at time of writing, the work represents a significant step toward scalable, observation-driven world models and has attracted attention within the robotics and embodied AI communities. Zholus's research places him at the intersection of computer vision, video representation learning, and robot learning — areas experiencing rapid growth as the field moves toward more autonomous, adaptable AI systems capable of acting meaningfully in unstructured real-world environments.
Research Focus
Key Achievements
Top Papers
- 1