Dylan Hadfield-Menell

Assistant Professor at MIT
PhD Student, Computer Science (2013-2020)
Dylan will be starting as an Assistant Professor at the Massachusetts Institute of Technology in July 2021.
While at CHAI, Dylan was a Ph.D. student at UC Berkeley advised by Anca Dragan, Pieter Abbeel, and Stuart Russell. His research focused on algorithms that facilitate human-compatible artificial intelligence. In particular, he tried to develop frameworks that account for uncertainty about the objective being optimized.
Before coming to Berkeley, Dylan did a Master’s of Engineering with Leslie Kaelbling and Tomás Lozano-Pérez at MIT. At Berkeley, Dylan’s research has taken a turn to focus more on AI safety and, thinking longer term, AI value alignment.
In 2016, he and his advisors formally described a cooperative inverse reinforcement learning problem (paper). The problem serves as a tool to help researchers consider how robots could learn humans’ values via cooperative instruction. While robots could learn humans’ values by observing humans, cooperative instruction is likely to be significantly faster.
In 2017, Dylan and his advisors described an “off-switch game,” a simplified problem describing scenarios in which a human would like to turn off a robot, but the robot is able to disable its off-switch (paper). They showed that a robot who is uncertain about the utility derived from various outcomes in the game is more likely to allow a human to turn it off.
You can learn more about Dylan’s other professional work at his website and follow him on Twitter here.
