Deepanway Ghosal

University of Michigan–Ann Arbor

Papers

1

Total Citations

4

H-Index

1

About

Deepanway Ghosal is a researcher advancing the frontiers of multimodal AI, with a primary focus on visual question answering (VQA) and language-guided vision-language models. His work addresses the critical challenge of enabling AI systems to not only "see" images but also reason about them in response to natural language queries. In his notable 2023 paper, "Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched Prompts," Ghosal introduced a novel approach that enriches prompts with external knowledge to significantly boost the performance of multimodal language models on VQA tasks. This contribution, which has garnered 4 citations in a short time, demonstrates his ability to bridge the gap between raw visual data and sophisticated reasoning. By integrating structured knowledge into the prompting process, Ghosal’s research paves the way for more context-aware and accurate AI systems, with potential applications in assistive technologies, automated content understanding, and human-computer interaction. His work stands out for its practical impact on making multimodal models more interpretable and effective, marking him as an emerging voice in the dynamic field of vision-language understanding.

Research Focus

Key Achievements

1
H-Index
1
Papers
4
Total Citations
4
Avg Citations/Paper
🏆 Most Cited Paper
Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched Prompts
4 citations · 2023
📈 Most Prolific Year: 2023 (1 Papers)
🤝 Key Collaborators: 4
🏛 Institutions: University of Michigan–Ann Arbor

Top Papers

  1. 1

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 12 days ago