Manli Zhu
Papers
1
Total Citations
3
H-Index
1
About
Manli Zhu is a rising researcher in artificial intelligence, whose work centers on advancing visual question answering (VQA) and multimodal reasoning systems. Her key contributions lie at the intersection of computer vision and natural language processing, particularly in enhancing the interpretability and precision of AI models. In her notable 2024 paper, "Detect2Interact: Localizing Object Key Field in Visual Question Answering with LLMs," Zhu introduces a novel framework that integrates large language models (LLMs) with fine-grained object localization. This work addresses a critical limitation in VQA—enabling systems to not only answer questions but also pinpoint and interact with specific parts of objects in images, thereby improving spatial accuracy and contextual relevance. Though early in her career, this paper has already garnered 3 citations, signaling its growing influence. Zhu’s research promises to bridge the gap between high-level reasoning and low-level visual perception, with potential applications in robotics, assistive technologies, and human-computer interaction. Her innovative approach to combining LLMs with localization techniques marks her as a promising contributor to the next generation of intelligent, context-aware AI systems.
Research Focus
Key Achievements
Top Papers
- 1