Daniel Filan

Research Manager at MATS

PhD Student in EECS

Daniel Filan is a PhD student of EECS at UC Berkeley, supervised by Stuart Russell. He’s interested in effective altruism and wants to ensure that future artificial intelligences who may be much more strategically intelligent than us behave in a safe way.

In 2016, Daniel worked with Tom Everitt, Mayank Daswani, and Marcus Hutter, considering the problem that a sufficiently advanced AI could choose to modify its source code in order to have easily achievable goals, and such modifications may not be to humans’ liking (paper). They determined that an agent will not self-modify if and only if the value function of the agent anticipates the consequences of self-modification and uses the agent’s current utility function when evaluating the future.

Daniel is currently thinking about mechanistic transparency, the problem of how to understand the workings of trained models.

You can learn more about Daniel’s work at his personal website.