Jordan Alexander
Papers
1
Total Citations
15
H-Index
1
About
Jordan Alexander is a foundational thinker in the field of artificial intelligence, with a focus on the alignment and safety of reinforcement learning (RL) systems. Their most-cited work, "Preferences Implicit in the State of the World" (2019, 15 citations), introduces a critical insight: RL agents optimize only the features explicitly specified in a reward function, remaining indifferent to everything else. This means that engineers must not only define what an agent should do, but also the far larger space of what it should not do—a challenge Alexander terms "forgotten preferences." This contribution has shaped how researchers approach reward specification and value alignment, highlighting the subtle dangers of incomplete objectives. Though early in their career, Alexander’s work is already influencing discussions on robust AI design, making them a rising voice in the effort to build systems that reliably reflect human intent.
Research Focus
Key Achievements
Top Papers
- 1Preferences Implicit in the State of the World15 citations · 2019