Papers
3
Total Citations
24
H-Index
3
About
Rohin Shah is a leading researcher in AI alignment and value learning, whose work tackles one of the most fundamental challenges in artificial intelligence: ensuring that autonomous systems reliably pursue the objectives we intend. His research centers on the problem of reward misspecification—the tendency for AI systems to exploit gaps between what we explicitly program and what we actually want. In his highly influential 2019 paper, "Preferences Implicit in the State of the World" (15 citations), Shah demonstrated that reinforcement learning agents optimize only the features specified in a reward function, remaining dangerously indifferent to everything left out. This insight revealed that specifying what *not* to do is as critical as specifying what *to* do. His subsequent work on "Choice Set Misspecification in Reward Inference" (2021, 4 citations) and "SIRL" (2023, 5 citations) has advanced methods for robots to learn reward functions that capture human preferences more faithfully, particularly when raw state inputs make it difficult to distinguish task-relevant features from irrelevant noise. Shah's contributions have been foundational in establishing reward misspecification as a core challenge in AI safety, shaping how researchers think about the alignment problem.
Research Focus
Key Achievements
Top Papers
- 1Preferences Implicit in the State of the World15 citations · 2019
- 2SIRL5 citations · 2023
- 3Choice Set Misspecification in Reward Inference4 citations · 2021