Papers
4
Total Citations
38
H-Index
3
About
Ziad Al-Halah is a rising star in artificial intelligence, whose research sits at the exciting intersection of computer vision, audio processing, and embodied AI. His primary focus is on creating agents that can perceive and interact with the world as humans do—by intelligently integrating what they see and hear. Al-Halah’s major contributions lie in pioneering the field of *active audio-visual learning*. In his highly-cited work "Move2Hear" (2021), he introduced the novel problem of an agent that must physically move to better separate a target sound from a noisy environment, a crucial step for real-world robotics. He further advanced this with "Few-Shot Audio-Visual Learning of Environment Acoustics" (2022, 17 citations), which enables agents to rapidly learn how a room’s geometry changes sound, with profound implications for AR and VR. His most recent work, "NaQ" (2023, 15 citations), tackles the challenge of searching through long egocentric video streams using natural language, effectively building a "memory" for AI agents. By tackling these fundamental problems, Al-Halah is not just publishing papers; he is defining the roadmap for the next generation of perceptually-aware, interactive machines.
Research Focus
Key Achievements
Top Papers
- 1Few-Shot Audio-Visual Learning of Environment Acoustics17 citations · 2022
- 2NaQ: Leveraging Narrations as Queries to Supervise Episodic Memory15 citations · 2023
- 3Move2Hear: Active Audio-Visual Source Separation4 citations · 2021
- 4DynGraph: Visual Question Answering via Dynamic Scene Graphs2 citations · 2019