Zeyuan Allen-Zhu

Massachusetts Institute of Technology

Papers

1

Total Citations

176

H-Index

1

About

Zeyuan Allen-Zhu is a leading theorist in machine learning, best known for his groundbreaking work on the optimization and generalization of deep neural networks. His research bridges deep learning theory, optimization algorithms, and large-scale machine learning, with a particular focus on understanding why over-parameterized networks succeed. His landmark 2018 paper, "A Convergence Theory for Deep Learning via Over-Parameterization" (176 citations), provided one of the first rigorous proofs that gradient descent can converge to global minima for deep networks, establishing a foundational framework for the field. Beyond this, he has made seminal contributions to variance reduction methods, including the SVRG and Katyusha algorithms, and to the theory of non-convex optimization, adversarial robustness, and the role of noise in training. His work has been recognized with multiple best paper awards and nominations at top venues like NeurIPS and ICML. With thousands of citations across his publications, Allen-Zhu’s research has profoundly shaped how the community understands the inner workings of deep learning, offering both elegant mathematical insights and practical implications for training modern neural networks.

Research Focus

Key Achievements

1
H-Index
1
Papers
176
Total Citations
176
Avg Citations/Paper
🏆 Most Cited Paper
A Convergence Theory for Deep Learning via Over-Parameterization
176 citations · 2018
📈 Most Prolific Year: 2018 (1 Papers)
🤝 Key Collaborators: 2
🏛 Institutions: Massachusetts Institute of Technology

Top Papers

  1. 1

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 11 days ago