Lawrence Chan

Member of Technical Staff at METR
PhD Student (2023), Advised by Anca Dragan and Stuart Russell
As of January 2023, Lawrence is working at METR (formerly “ARC Evals”), where he conducts evaluations of large language models. Previously, he was at Redwood Research, focusing on adversarial training and neural network interpretability.
He is also pursuing a PhD at UC Berkeley under the supervision of Anca Dragan and Stuart Russell. Before that, he earned a BAS in Computer Science and Logic and a BS in Economics from the University of Pennsylvania’s M&T Program, where he had the opportunity to work with Philip Tetlock on applying machine learning to forecasting.
Lawrence’s primary research interests include mechanistic interpretability and scalable oversight. He has also explored conceptual questions related to learning human values.
Additionally, he occasionally blogs about AI alignment and other topics on LessWrong and the AI Alignment Forum.
