Papers

1

Total Citations

4

H-Index

1

About

Linfeng Xu is a leading researcher in multimodal perception and audio-visual machine learning, with a particular focus on bridging the gap between auditory and visual modalities for scene understanding. His most influential work introduces the concept of cross-modal cognitive consensus for audio-visual segmentation (AVS), a pioneering framework that enables precise pixel-wise extraction of sounding objects from video frames. This breakthrough, published in 2024, has already garnered 4 citations and addresses critical challenges in multi-modal video editing, augmented reality, and intelligent robotic systems. Xu’s contributions establish a foundational paradigm for how machines can jointly interpret sound and sight, moving beyond simple object detection to achieve fine-grained, semantically-aware segmentation guided by audio cues. His research has significant implications for developing more intuitive human-computer interfaces and autonomous systems that can perceive their environment as holistically as humans do. By formalizing the AVS task and demonstrating its feasibility, Xu has opened new avenues for research in embodied AI and multi-modal reasoning, positioning him as a key innovator in this rapidly evolving field.

Research Focus

Key Achievements

1
H-Index
1
Papers
4
Total Citations
4
Avg Citations/Paper
🏆 Most Cited Paper
Cross-Modal Cognitive Consensus Guided Audio–Visual Segmentation
4 citations · 2024
📈 Most Prolific Year: 2024 (1 Papers)
🤝 Key Collaborators: 4
🏛 Institutions: University of Electronic Science and Technology of China

Top Papers

  1. 1

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 11 days ago