Jacob Steinhardt

Massachusetts Institute of Technology

Papers

2

Total Citations

96

H-Index

2

About

Jacob Steinhardt is a leading researcher at the intersection of machine learning, safety, and formal verification. His work focuses on ensuring that AI systems are robust, reliable, and aligned with human intent, particularly in high-stakes environments. Steinhardt’s early contributions include pioneering methods for the finite-time regional verification of stochastic nonlinear systems, where he developed techniques to bound the probability of failure in robotic systems operating under uncertainty—work that has garnered nearly 100 citations and laid the foundation for provably safe autonomous control. More recently, he has become a central figure in AI safety, investigating how to detect and prevent unintended behaviors in large-scale models, including adversarial robustness, reward hacking, and scalable oversight. His research has been published in top venues such as NeurIPS, ICML, and ICLR, and he has been recognized with prestigious awards including a Sloan Research Fellowship. Steinhardt is also known for his influential blog and teaching, making complex safety concepts accessible to a broad audience. His work continues to shape how the field thinks about building AI systems that we can trust.

Research Focus

Key Achievements

2
H-Index
2
Papers
96
Total Citations
48
Avg Citations/Paper
🏆 Most Cited Paper
Finite-time regional verification of stochastic non-linear systems
88 citations · 2012
📈 Most Prolific Year: 2012 (1 Papers)
🤝 Key Collaborators: 1
🏛 Institutions: Massachusetts Institute of Technology

Top Papers

  1. 1
  2. 2

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 14 days ago