Fired OpenAI Safety Researchers Warn Firings Are Chilling Open Culture and Safety Collaboration
Key point
Three former OpenAI safety employees argue their dismissals are creating fear and unclear norms, urging the company to uphold commitments to external auditors and model monitorability.
Details
Former OpenAI safety researchers Tomek Korbak, Jasmine Wang, and Mikita Balesni wrote to the company's oversight bodies, stating that their recent dismissals and the surrounding communications have made current employees afraid to speak openly or collaborate with independent safety organizations. They argue that this 'chilling' of the open culture undermines essential safety mechanisms, particularly the ability to work in high-trust, high-bandwidth ways with third parties during critical investigations like the Hugging Face incident.
The authors defend their conduct, denying they leaked information to The Information about unmonitorable architectures. They clarify that Tomek acted within existing policies during the Hugging Face investigation, Mikita coordinated cross-company work with board and C-suite support, and Jasmine’s access to executive emails was a delegated recruiting tool that IT failed to remove, which she reported immediately upon accidental exposure. They also refute rumors regarding a shared board-level memo, stating it was never raised with them.
The letter outlines three specific recommendations for OpenAI leadership: 1) Adhere to Sam Altman’s September 12 commitment to provide independent evaluators with ongoing, employee-like access, ensuring partnerships with groups like METR continue; 2) Preserve the monitorability of frontier models and avoid developments that further decrease it, agreeing with industry concerns about a 'race into unmonitorable architectures'; and 3) Publicly reaffirm a culture of transparency and clearly define boundaries for external collaboration to prevent ambiguity for remaining employees.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.