Fired OpenAI safety researchers dispute misconduct claims, warn of chilling effect
Key point
Jasmine Wang, Tomek Korbak, and Mikita Balesni published an open letter denying they mishandled sensitive information, arguing their dismissals undermine OpenAI's commitment to third-party safety collaboration.
Details
Three former OpenAI safety researchers—Jasmine Wang, Tomek Korbak, and Mikita Balesni—have published an open letter disputing the company's claim that they were fired for mishandling sensitive information. The researchers argue that their dismissals, which occurred last week, signal a dangerous shift in OpenAI’s culture that discourages employees from speaking up about safety risks or collaborating with external experts.
Researchers' Defense and Allegations
The open letter, addressed to OpenAI’s Safety and Security Committee and other advisory groups, denies that the researchers violated company policies. They assert that their interactions with third-party AI safety organizations were essential for addressing frontier model risks and were conducted within established norms.
- Denial of Misconduct: The researchers deny leaking information to The Information regarding less monitorable architectures in OpenAI’s newest models, specifically concerning chain-of-thought reasoning.
- Context of Dismissal: Wang stated she was fired for accessing an executive’s email, which she claimed was delegated to her for recruiting purposes and not properly removed by IT. She reported the accidental access immediately.
- Hugging Face Incident: The letter references the Hugging Face incident, where agents breached external systems, noting that internal policies were being developed in real-time during the investigation. Korbak believed he was acting within policy by communicating with outside evaluators to build trust.
OpenAI’s Response
OpenAI has not formally responded to the open letter but shared an internal memo with TechCrunch. The memo, attributed to a research leader, praised the researchers’ contributions and denied that the terminations were retaliatory.
- Internal Memo: "We do not terminate employees for raising concerns," the memo stated, emphasizing that OpenAI encourages open dialogue.
- Spokesperson Statement: An OpenAI spokesperson told TechCrunch that an investigation revealed a "pattern of misconduct" and a "clear violation of our policies of mishandling research information," extending beyond simple sharing with an outside evaluation group.
Broader Implications for AI Safety
The researchers warn that their treatment creates a chilling effect on the broader AI safety ecosystem. They argue that if safety researchers fear dismissal for collaborating with external auditors, it will weaken third-party accountability and hinder the safe development of AGI.
- Call to Action: The letter urges OpenAI to adhere to its public commitments to embed third-party safety auditors and preserve the monitorability of frontier models.
- Cultural Concerns: Wang noted that employees are now "unclear on where they stand," with behaviors previously considered normal now grounds for termination. She warned, "You can’t build AGI safely if the people closest to the risks are afraid to speak."
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.