Papers
9
Total Citations
76
H-Index
5
About
Jiafei Duan is an emerging robotics and AI researcher whose work sits at the intersection of physical reasoning, robot manipulation, and multimodal large language models. His research addresses a fundamental challenge in modern robotics: enabling robots to understand, reason about, and interact with the physical world in open-ended, real-world settings. Among his most notable contributions is **The Colosseum**, a benchmark for evaluating generalization in robotic manipulation (22 citations), and **Octopi**, a tactile-language model that incorporates touch as a modality for physical reasoning (21 citations) — a novel direction that extends beyond the dominant vision-language paradigm. His **NEWTON** framework (12 citations) critically examined whether large language models truly possess physical reasoning capabilities, sparking important conversations in the community. Further work includes **RoboPoint**, addressing spatial affordance prediction, and **AHA**, a system for detecting and learning from robotic failures — both reflecting his commitment to robust, self-improving robotic systems. Duan also explores democratizing robot training through augmented reality (**EVE**) and automating real-world manipulation via vision-language models (**Manipulate-Anything**). With publications spanning benchmarking, multimodal reasoning, and human-robot interaction, his growing citation record signals rising influence in embodied AI research.
Research Focus
Key Achievements
Top Papers
- 1
- 2Octopi: Object Property Reasoning with Large Tactile-Language Models21 citations · 2024
- 3NEWTON: Are Large Language Models Capable of Physical Reasoning?12 citations · 2023
- 4
- 5
- 6
- 7EVE: Enabling Anyone to Train Robots using Augmented Reality3 citations · 2024
- 8Octopi: Object Property Reasoning with Large Tactile-Language Models2 citations · 2024
- 9