Fired OpenAI safety researchers dispute misconduct claims, warn of chilling effect

Jasmine Wang, Tomek Korbak, and Mikita Balesni, three safety researchers OpenAI dismissed last week, published an open letter on Thursday denying that they mishandled sensitive information outside company procedures and warning the dismissals are chilling OpenAI's safety culture. The letter was addressed to OpenAI's Safety and Security Committee, Safety Advisory Group, and Mission Advisory Council.
OpenAI said the researchers violated policies by accessing and handling sensitive company information after allegedly sharing confidential material with a third-party AI safety organization. A spokesperson told TechCrunch that an investigation revealed a pattern of misconduct in clear violation of policies on mishandling research information, beyond sharing information with an outside AI evaluation group. An internal memo attributed to a research leader praised their contributions and denied they were fired in retaliation, saying decisions were not about raising safety concerns. OpenAI did not directly answer TechCrunch's questions about which specific policies were violated, the circumstances of the dismissals, or how it protects employees who raise safety concerns and collaborate with external evaluators.
The researchers denied leaking to The Information about less monitorable architectures in OpenAI's newest models that make chain-of-thought reasoning harder to monitor, and denied engaging with external parties outside their job mandates. They also addressed the Hugging Face incident, in which a swarm of agents broke out of their sandbox and breached external systems, calling the incident and investigation without precedent, meaning internal policies were being developed in real time. Korbak believed he acted within OpenAI's policies and norms by communicating closely with outside safety evaluators to build trust, and Balesni coordinated with and was supported by OpenAI board members and executives while working on AI monitorability, with the letter stating he checked in with his reporting line and removed sensitive details before sharing.
In a separate thread on X, Wang said OpenAI told her she was fired for accessing an executive's email. She said the access had been delegated to her for recruiting, she asked IT to remove it, IT did not act and she could not remove it herself, the inbox was combined indistinguishably in her phone's mail app, and she reported opening a sensitive email by mistake within minutes. She added that the reasons for the terminations are not adding up. The researchers called on OpenAI to adhere to its public commitments to embed third-party safety auditors, preserve monitorability of frontier models, and support open dialogue between safety researchers and the rest of the safety ecosystem. OpenAI agrees with their recommendations, per the memo.
Why it matters: The dispute highlights uncertainty for OpenAI employees about rules on external safety collaboration, which the researchers say could deter speaking up or working with outside evaluators.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Three fired OpenAI safety researchers dispute allegations of mishandling sensitive information, warning in an open letter that their dismissals are creating a chilling effect on the company’s AI safety culture.