Cheng Xian
Papers
1
Total Citations
1
H-Index
1
About
Cheng Xian is a researcher at the forefront of multimodal learning and vision-language understanding, with a particular focus on composed image retrieval—a challenging task that combines visual and textual queries to refine search results. Their most notable work, "Fine-tuning CLIP for difference-guided composed image retrieval" (2025), introduces a novel approach to leveraging large-scale pre-trained models like CLIP for more precise, user-guided image searches. By fine-tuning CLIP to interpret "difference" instructions—such as "find an image similar to this one but with a different color"—Xian’s method significantly improves retrieval accuracy and interpretability, bridging the gap between natural language and visual semantics. While still early in its citation impact, this contribution addresses a critical bottleneck in interactive AI systems, with potential applications in e-commerce, design, and content management. Xian’s research exemplifies how targeted fine-tuning of foundation models can unlock new capabilities, making them more responsive to nuanced human intent. Their work is a valuable resource for students and researchers exploring vision-language models, retrieval systems, and human-AI interaction.
Research Focus
Key Achievements
Top Papers
- 1Fine-tuning CLIP for difference-guided composed image retrieval1 citations · 2025