Papers
1
Total Citations
4
H-Index
1
About
Qingbo Wu is a leading researcher in multimodal perception and computer vision, with a focus on audio-visual segmentation (AVS)—a task that extracts sounding objects from video frames to enable applications like multi-modal video editing, augmented reality, and intelligent robotics. His pioneering work, "Cross-Modal Cognitive Consensus Guided Audio–Visual Segmentation" (2024), introduces a novel framework that leverages cognitive consensus between audio and visual modalities to achieve precise pixel-wise segmentation, addressing a critical gap in cross-modal understanding. With over 4 citations in its first year, this paper has quickly become a foundational reference in the field. Wu’s contributions extend to advancing how machines integrate auditory and visual cues, pushing the boundaries of human-like perception in AI systems. His research is widely recognized for its impact on intelligent robot systems and augmented reality, where accurate audio-visual alignment is essential. Through his innovative approaches, Wu continues to shape the future of multimodal learning, inspiring new directions in interactive AI and scene understanding.
Research Focus
Key Achievements
Top Papers
- 1Cross-Modal Cognitive Consensus Guided Audio–Visual Segmentation4 citations · 2024