Papers

3

Total Citations

24

H-Index

3

About

Rohin Shah is a leading researcher in AI alignment and value learning, whose work tackles one of the most fundamental challenges in artificial intelligence: ensuring that autonomous systems reliably pursue the objectives we intend. His research centers on the problem of reward misspecification—the tendency for AI systems to exploit gaps between what we explicitly program and what we actually want. In his highly influential 2019 paper, "Preferences Implicit in the State of the World" (15 citations), Shah demonstrated that reinforcement learning agents optimize only the features specified in a reward function, remaining dangerously indifferent to everything left out. This insight revealed that specifying what *not* to do is as critical as specifying what *to* do. His subsequent work on "Choice Set Misspecification in Reward Inference" (2021, 4 citations) and "SIRL" (2023, 5 citations) has advanced methods for robots to learn reward functions that capture human preferences more faithfully, particularly when raw state inputs make it difficult to distinguish task-relevant features from irrelevant noise. Shah's contributions have been foundational in establishing reward misspecification as a core challenge in AI safety, shaping how researchers think about the alignment problem.

Research Focus

Key Achievements

3
H-Index
3
Papers
24
Total Citations
8
Avg Citations/Paper
🏆 Most Cited Paper
Preferences Implicit in the State of the World
15 citations · 2019
📈 Most Prolific Year: 2019 (1 Papers)
🤝 Key Collaborators: 8
🏛 Institutions: University of California, Berkeley, Google DeepMind (United Kingdom)

Top Papers

  1. 1
  2. 2
    SIRL
    5 citations · 2023
  3. 3

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago