OpenAI’s safety oversight will now include Paul Christiano, an alignment researcher who says rapid advances in AI capabilities could create a meaningful risk of catastrophic, irreversible loss of human control. Christiano is joining the OpenAI Foundation Board and its Safety and Security Committee, while serving as a non-voting observer on the company’s for-profit board.
An oversight role across the company
The Foundation’s Safety and Security Committee provides governance over safety and security practices across OpenAI, including OpenAI Group PBC. Christiano will serve alongside committee chair Zico Kolter. The Foundation itself controls OpenAI Group PBC while operating as a separate charitable organization, according to OpenAI.
OpenAI said Christiano brings experience from government and alignment research. He is a senior technical adviser at the Center for AI Standards and Innovation within NIST and founded the Alignment Research Center. He previously led alignment research at OpenAI from 2017 to 2021. As a NIST adviser, OpenAI said, he will recuse himself from OpenAI-related matters and model evaluations.
AI capabilities have advanced very rapidly in the last year and alignment remains a difficult technical problem, making the Safety and Security Committee’s responsibility more important and more challenging than ever.
Paul Christiano, in OpenAI’s announcement
A warning, not an endorsement
In his personal statement, Christiano said he does not believe the AI industry, including OpenAI, is currently on track to reduce loss-of-control risk to an acceptable level. He also said joining OpenAI should not be read as either an endorsement or a criticism of its safety practices in particular.
His concern centers on automated AI research and development. Christiano argued that, if AI systems could fully automate that work, better training and algorithms could increase the number and quality of AI researchers doing further research. He said that feedback loop could potentially overcome diminishing returns and computing limits, producing rapid capability gains. This is a forecast, not a claim that such automation has already occurred.
Christiano cited OpenAI’s prediction that systems might gain enough capability to fully automate AI research within 18 months. But he called his own timing forecast highly uncertain, putting it anywhere from several months to several years. He said a failure to develop more robust alignment before superintelligent systems emerge could permanently remove human control.
The standard he wants developers to meet
- Improve safety mitigations, including slowing development when necessary.
- Share evidence about risks and the effectiveness of mitigations transparently.
- Work toward shared safety standards and domestic and international coordination.
Christiano’s appointment gives those arguments a formal place in OpenAI’s oversight structure, but it does not itself settle whether the company’s safeguards are sufficient. His stated benchmark is that OpenAI and other developers should be judged by externally verifiable behavior and results. That puts the emphasis on what governance delivers, rather than on who occupies a seat.
Reader comments
Newest comments first. Replies stay oldest first.