Hirokatsu Kataoka
National Institute of Advanced Industrial Science and Technology
Papers
5
Total Citations
67
H-Index
4
About
Hirokatsu Kataoka is a leading researcher at the forefront of computer vision and human-robot interaction, specializing in scene understanding, visual question answering (VQA), and change captioning. His major contributions lie in developing frameworks that enable machines to perceive and describe dynamic 3D environments from multiple viewpoints. Kataoka pioneered the concept of "scene change captioning," creating systems that recognize and articulate changes in indoor and real-world scenes using natural language—a critical capability for applications like anomaly detection and autonomous robotics. His work on multi-view VQA, which allows iterative observation to answer spatial questions, advances how AI handles geometric information beyond single-image inputs. With over 67 citations across his most-cited papers, Kataoka’s research bridges vision and language, notably in his 2020 letter on 3D-aware change captioning (26 citations) and his 2020 study on multimodal indoor scene changes (23 citations). His innovative integration of 3D data into VQA tasks marks a significant step toward more context-aware, interactive AI systems, making his work essential for students and researchers exploring embodied intelligence and scene reasoning.
Research Focus
Key Achievements
Top Papers
- 13D-Aware Scene Change Captioning From Multiview Images26 citations · 2020
- 2Indoor Scene Change Captioning Based on Multimodality Data23 citations · 2020
- 3Multi-View Visual Question Answering with Active Viewpoint Selection12 citations · 2020
- 4Incorporating 3D Information Into Visual Question Answering4 citations · 2019
- 5Scene Change Captioning in Real Scenarios2 citations · 2022