Paul Christiano
American AI safety researcher
From Wikipedia, the free encyclopedia
Paul Christiano is an American researcher in the field of artificial intelligence (AI), with a specific focus on AI alignment, which is the subfield of AI safety research that aims to steer AI systems toward human interests.[1] He is the founder and executive director of the Alignment Research Center (ARC), a nonprofit that develops methods for finding mechanistic explanations of neural network behavior,[2] and serves on the Safety and Security Committee of the OpenAI Foundation's board.[3]
Paul Christiano | |
|---|---|
| Education | |
| Known for | |
| Scientific career | |
| Workplaces | |
| Thesis | Manipulation-resistant online learning (2017) |
| Umesh Vazirani | |
| Website | paulfchristiano |
Christiano worked at OpenAI from 2017 to 2021, where he led the language model alignment team[4] and became one of the principal architects of reinforcement learning from human feedback (RLHF).[5][6] He founded ARC in 2021[7] and in 2024 became Head of Safety for the U.S. AI Safety Institute (now the Center for AI Standards and Innovation) inside NIST,[8] before returning to ARC as executive director.[2] He was also an initial trustee of Anthropic's Long-Term Benefit Trust.[6][9] In 2023, Christiano was named to the TIME 100 Most Influential People in AI[5][10] and appointed to the advisory board of the UK government's Frontier AI Taskforce.[11]
Education
Christiano attended the Harker School in San Jose, California.[12] He competed on the U.S. team and won a silver medal at the 49th International Mathematics Olympiad (IMO) in 2008.[12][13]
In 2012, Christiano graduated from the Massachusetts Institute of Technology (MIT) with a degree in mathematics.[14][15] At MIT, he researched data structures, quantum cryptography, and combinatorial optimization.[15]
He then went on to complete a PhD at the University of California, Berkeley.[16] While at Berkeley, Christiano collaborated with researcher Katja Grace on AI Impacts, co-developing a preliminary methodology for comparing supercomputers to brains, using traversed edges per second (TEPS).[17] He also experimented with putting Carl Shulman's donor lottery theory into practice, raising nearly $50,000 in a pool to be donated to a single charity.[18]
Career
At OpenAI, Christiano co-authored the paper "Deep Reinforcement Learning from Human Preferences" (2017) and other works developing reinforcement learning from human feedback (RLHF).[19][20] He is considered one of the principal architects of RLHF,[5][6] which in 2017 was "considered a notable step forward in AI safety research", according to The New York Times.[21] Other works such as "AI safety via debate" (2018) focus on the problem of scalable oversight – supervising AIs in domains where humans would have difficulty judging output quality.[22][23][24]
Christiano left OpenAI in 2021 to work on more conceptual and theoretical issues in AI alignment and subsequently founded the Alignment Research Center to focus on this area.[1] One subject of study is the problem of eliciting latent knowledge from advanced machine learning models.[25][26] ARC also develops techniques to identify and test whether an AI model is potentially dangerous.[5] In April 2023, Christiano told The Economist that ARC was considering developing an industry standard for AI safety.[27]
As of April 2024, Christiano was listed as the head of AI safety for the US AI Safety Institute at NIST.[28] One month earlier in March 2024, staff members and scientists at the institute threatened to resign upon being informed of Christiano's pending appointment to the role, stating that his ties to the effective altruism movement may jeopardize the AI Safety Institute's objectivity and integrity.[29]
Views on AI risks
He is known for his views on the potential risks of advanced AI. In 2017, Wired magazine stated that Christiano and his colleagues at OpenAI weren't worried about the destruction of the human race by "evil robots", explaining that "[t]hey’re more concerned that, as AI progresses beyond human comprehension, the technology’s behavior may diverge from our intended goals."[30]
In a widely quoted interview with Business Insider in 2023, Christiano said that there is a “10–20% chance of AI takeover, [with] many [or] most humans dead.” He also conjectured a “50/50 chance of doom shortly after you have AI systems that are human level.”[31][1]