Papers
5
Total Citations
68
H-Index
4
About
Kenji Iwata is a leading researcher in multimodal scene understanding, with a focus on bridging the gap between 3D spatial perception and natural language. His core contributions lie in developing frameworks that enable robots and AI systems to recognize, describe, and reason about changes in real-world environments. Iwata pioneered the concept of scene change captioning, creating systems that generate natural language descriptions of alterations observed in indoor and multi-view 3D scenes—a critical capability for human-robot interaction and anomaly detection. His most cited work, "3D-Aware Scene Change Captioning From Multiview Images" (26 citations), demonstrates how to synthesize observations from multiple viewpoints into coherent text. He has also advanced visual question answering by introducing active viewpoint selection, allowing agents to iteratively explore scenes to answer queries. Notably, his early work on the musician robot’s speech conversation system (1985) shows a long-standing commitment to embodied AI. With over 68 citations across his top papers, Iwata’s research continues to shape how machines perceive, describe, and interact with dynamic 3D environments.
Research Focus
Key Achievements
Top Papers
- 13D-Aware Scene Change Captioning From Multiview Images26 citations · 2020
- 2Indoor Scene Change Captioning Based on Multimodality Data23 citations · 2020
- 3Multi-View Visual Question Answering with Active Viewpoint Selection12 citations · 2020
- 4SPEECH CONVERSATION SYSTEM OF THE MUSICIAN ROBOT.5 citations · 1985
- 5Scene Change Captioning in Real Scenarios2 citations · 2022