Yuxuan Bai

Xi'an Polytechnic University

Papers

1

Total Citations

18

H-Index

1

About

Yuxuan Bai is a rising researcher in artificial intelligence, with a primary focus on multimodal generative models and human-computer interaction. Bai’s most cited work, “Integrated visual transformer and flash attention for lip-to-speech generation GAN” (2024), addresses a critical challenge in lip-to-speech (LTS) synthesis—an emerging technology with applications ranging from assisting speech-impaired individuals to enhancing virtual assistants and robots. By integrating visual transformers with flash attention mechanisms into a generative adversarial network framework, Bai’s approach improves the fidelity and efficiency of reconstructing intelligible speech from silent video. This contribution has already garnered 18 citations, signaling strong early impact in a rapidly evolving field. Bai’s research sits at the intersection of computer vision, natural language processing, and assistive technology, pushing the boundaries of how machines can interpret and generate human-like communication. As LTS continues to mature, Bai’s work provides a foundational step toward more natural, real-time speech interfaces. For students and researchers exploring multimodal AI, Bai’s research offers a compelling example of how attention-based architectures can bridge visual and auditory modalities for real-world benefit.

Research Focus

Key Achievements

1
H-Index
1
Papers
18
Total Citations
18
Avg Citations/Paper
🏆 Most Cited Paper
Integrated visual transformer and flash attention for lip-to-speech generation GAN
18 citations · 2024
📈 Most Prolific Year: 2024 (1 Papers)
🤝 Key Collaborators: 3
🏛 Institutions: Xi'an Polytechnic University

Top Papers

  1. 1

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 12 days ago