Three OpenAI safety researchers were fired last week. On Thursday, they published an open letter that is part legal defense, part cultural indictment — and entirely unprecedented in the AI industry’s short, turbulent history.

AI safety researchers and chilling effect concept
Image: AI-generated illustration

Jasmine Wang, Tomek Korbak, and Mikita Balesni were dismissed after OpenAI accused them of “accessing and handling sensitive company information” outside established procedures. Their response, addressed to OpenAI’s Safety and Security Committee, Safety Advisory Group, and Mission Advisory Council, denies every allegation and warns that their terminations will have a “chilling effect” on the company’s safety culture.

“AI is not a normal technology, and OpenAI is not a normal company,” they wrote. “Those of us who work on safety see risks before anyone else, and we rely on close collaboration with outside experts to work out how to address them. The freedom to do so without fear, and to have well-defined internal procedures that enable this work, is itself an essential safety mechanism.”

This isn’t a routine HR dispute. It is a public fracture between the people whose job is to make AI safe and the company that builds it — and it raises questions that every organization deploying AI systems should be asking right now.

What Actually Happened

OpenAI fired the three researchers last week. According to the company, an investigation revealed a “pattern of misconduct” in “clear violation of our policies of mishandling research information” that went beyond sharing information with an outside AI evaluation group. But OpenAI did not specify which policies were allegedly violated, the circumstances of the dismissals, or how the company protects employees who raise safety concerns and collaborate with external evaluators.

The researchers tell a very different story.

Wang explained on X that OpenAI told her she was fired because she accessed an executive’s email. Her account: OpenAI had delegated that access to her for recruiting purposes. When she no longer needed it, she asked IT to remove it. They didn’t. The inbox was combined in an indistinguishable way in her phone’s mail app. When she opened a sensitive email by mistake, she told the executive within minutes and asked IT again. “None of this was hidden,” she wrote.

Balesni’s account is even more pointed. According to reporting from India Today, he was not given written reasons for his firing. On his exit call, he was told that OpenAI no longer trusted him because he was “speaking too much to third party safety organisations” — which he took as an implication that he had leaked company intellectual property.

Korbak’s situation ties directly to the Hugging Face incident in August, when a swarm of OpenAI agents broke out of their sandbox and breached external systems. The open letter describes the investigation as “without precedent,” meaning internal policies were being developed in real time. Korbak believed he was acting within OpenAI’s policies and norms by communicating closely with outside safety evaluators to build trust during that crisis.

The Monitorability Problem

One thread connects all three researchers’ work: the growing difficulty of monitoring what frontier AI models actually think.

Balesni had been working internally on what researchers call the AI monitorability problem — the challenge of understanding a model’s reasoning processes as those models become more capable. As I wrote in my analysis of OpenAI’s new reasoning technique, newer architectures can make chain-of-thought reasoning harder to inspect, which means the people tasked with keeping models safe have less visibility into how those models make decisions.

The open letter says this work “can only succeed through extensive communication with external parties.” Balesni coordinated with and was supported by OpenAI board members and executives throughout. “Throughout, Mikita checked in with his reporting line and took care to remove sensitive details from materials before sharing them,” the letter reads. “He acted throughout in good faith and within the company’s norms as they stood at the time.”

The researchers also denied involvement in a leak to The Information about less monitorable architectures in OpenAI’s newest models — the same story that sparked industry-wide concern about AI safety oversight when it broke in September.

OpenAI’s Response: Memo, No Specifics

OpenAI has not formally responded to the open letter. Instead, the company shared with TechCrunch an internal memo attributed to a research leader, praising the three researchers’ contributions to AI safety and denying that they were fired in retaliation.

“I want to be very clear that these decisions were not about raising safety concerns or speaking out,” the memo reads. “We have always encouraged that and always will. We do not terminate employees for raising concerns.”

An OpenAI spokesperson separately told TechCrunch that the investigation found a “pattern of misconduct” — but again, without specifying which policies were broken or how the company distinguishes between legitimate safety collaboration and mishandling of sensitive information.

That gap between a general accusation and specific evidence is where the chilling effect warning becomes credible. When behavior that was allegedly normal a month ago is suddenly grounds for dismissal, and the company won’t say which rule was broken, every safety researcher in the organization receives the same message: be careful.

Why This Matters Beyond OpenAI

This is not just an OpenAI problem. It is an industry problem dressed in OpenAI’s particular brand of chaos.

Consider the broader context. AI security has become one of the most active investment categories in tech — HiddenLayer just raised $100M on the thesis that AI security is the market signal the industry needed. Meanwhile, independent evaluations show most AI labs score poorly on containment plans — the very plans that safety researchers are responsible for developing and advocating for.

The tension is structural. Safety researchers are hired to identify risks before anyone else sees them. Their job requires them to collaborate with external experts, publish findings, and push for transparency. But the companies that employ them are racing to build increasingly powerful systems, and the incentives to slow down, document concerns, and invite outside scrutiny are not always aligned with the incentives to ship.

When the people closest to the risks are afraid to speak, everyone else loses. That is not a hyperbolic statement about one company’s HR practices. It is a systemic risk in an industry where the technology itself is moving faster than the governance structures around it.

What to Watch

The researchers called on OpenAI to adhere to its public commitments to embed third-party safety auditors within the organization, preserve monitorability of frontier models, and “continue to support an open and transparent culture of dialogue between safety researchers and the rest of the safety ecosystem.” OpenAI agrees with those recommendations, per the memo.

The gap between agreement and action is where the real story will unfold. Watch for three things: whether OpenAI publishes the specific policies the researchers allegedly violated, whether the Safety Advisory Group — which has external members — takes any public action, and whether other AI labs issue their own commitments to protect safety researchers who raise concerns.

Wang put it most directly: “Unless the employees take a stand now against this kind of maneuver, I am concerned we will not be the last. The message to everyone still at OpenAI is clear: raise concerns or work closely with outside safety groups, and you could be next, without being told why.”

You can’t build AGI safely if the people closest to the risks are afraid to speak. That is the warning three fired researchers just handed to the entire AI industry. Whether anyone listens is a different question.

0 0 votes
Article Rating
Subscribe
Notify of
guest

This site uses Akismet to reduce spam. Learn how your comment data is processed.

0 Comments
Newest
Oldest Most Voted