Papers
1
Total Citations
2
H-Index
1
About
Cunhan Guo is a rising researcher in the field of multi-modal perception, with a primary focus on audio-visual semantic segmentation (AVSS) and its applications in robotic navigation and autonomous driving. His most notable contribution, the paper "ALOHA: Adapting Local Spatio-Temporal Context to Enhance the Audio-Visual Semantic Segmentation" (2025), addresses a critical limitation in existing AVSS methods that rely on global spatio-temporal modules for fusing audio and visual representations. By introducing a novel approach that adapts local spatio-temporal context, Guo’s work enhances pixel-level multi-modal understanding, enabling more precise and context-aware segmentation in dynamic environments. Although early in his career, with 2 citations to date, his research demonstrates significant potential for advancing real-world autonomous systems. Guo’s work is particularly impactful for students and researchers exploring the intersection of computer vision and audio processing, offering a fresh perspective on how localized information can improve semantic segmentation tasks. His contributions underscore the importance of fine-grained multi-modal fusion, positioning him as a promising voice in the development of more robust and efficient perception systems for next-generation robotics and autonomous vehicles.
Research Focus
Key Achievements
Top Papers
- 1