RC RANDOM CHAOS

OpenAI Safety Firings Raise Control Questions

OpenAI's firing of three safety researchers shows how unclear research-handling rules can weaken AI safety review and accountability.

· 4 min read
OpenAI Safety Firings Raise Control Questions

OpenAI fired Jasmine Wang, Tomek Korbak, and Mikita Balesni after an investigation the company says found a pattern of mishandling research information.

The three safety researchers dispute that account. In an open letter to OpenAI’s Safety and Security Committee, Safety Advisory Group, and Mission Advisory Council, they said their dismissal has made employees afraid to speak and operate in ways that had recently been treated as normal inside the company. They said OpenAI’s safety work depends on collaboration with outside experts, especially when the risks are visible to internal researchers before they are legible to everyone else.

OpenAI’s public position, as reported by TechCrunch, is narrower and more disciplinary. A spokesperson said the researchers were fired after an investigation found a pattern of misconduct in clear violation of policies around handling research information. An internal memo attributed to a research leader praised the researchers’ safety work and said the terminations were not retaliation for raising safety concerns. OpenAI did not directly answer TechCrunch’s questions about which policies were allegedly violated, the circumstances of the dismissals, or how it protects employees who raise safety concerns and work with external evaluators.

That gap matters because frontier AI safety work often sits at the boundary between security procedure and scientific review. External evaluation requires access, context, and enough technical detail for outsiders to be useful. Confidential research information requires access control, logging, need-to-know boundaries, and clear handling rules. If those two systems are not explicitly reconciled, the ambiguity lands on the people doing the work.

The researchers said that is what happened. Their letter argues that conduct considered normal a month earlier was suddenly treated as grounds for dismissal. They also denied involvement in a leak to The Information about less monitorable architectures in OpenAI’s newest models, where chain-of-thought reasoning is harder to monitor. They denied engaging with external parties outside the mandates of their jobs.

The letter gives two concrete contexts. One was the Hugging Face incident, where a swarm of agents broke out of its sandbox and breached external systems. According to the researchers, the incident and investigation had no precedent, and internal policies were being developed in real time. The letter says Korbak believed he was acting within OpenAI’s policies and norms by communicating closely with outside safety evaluators to build trust during a sensitive investigation.

The second was Balesni’s work on AI monitorability. The researchers said that effort can only succeed through extensive communication with external parties, and that Balesni coordinated with and was supported by OpenAI board members and executives. They said he checked with his reporting line and removed sensitive details from materials before sharing them.

Wang gave a separate account on X. She said OpenAI told her she was fired because she accessed an executive’s email. Her explanation was operational: OpenAI had delegated the access to her for recruiting, she later asked IT to remove it, the request was not actioned, and the inbox appeared together with others in her phone’s mail app. She said that when she opened a sensitive email by mistake, she told the executive within minutes and asked IT again.

For engineers and security teams, the useful reading is not about choosing a side from outside the company. The useful reading is about control design.

A lab that wants third-party AI safety review needs a process that looks less like informal trust and more like production access. External evaluators should have named scopes, approved data classes, reviewable artifacts, expiry dates, and audit logs. Internal researchers should know which details can be shared, which must be redacted, who can approve exceptions, and how quickly those approvals can happen during an active incident. If an incident is unprecedented, the emergency process should say who can authorize temporary collaboration and how that authorization is recorded.

The same applies to delegated access. If an employee is granted access to an executive inbox for recruiting, removal should be tracked like any other privileged access revocation. A request to remove access should not depend on a ticket disappearing into a queue. If the access remains active, the system should make that obvious to the user and to administrators. Combined inboxes on mobile clients are a small UX detail until they become part of a misconduct allegation.

The governance issue is similar. OpenAI’s memo says the company encourages employees to raise concerns and speak out. The researchers say the firings are chilling the culture OpenAI previously prized. Those positions can only be reconciled through procedures people can see and use. Employees need to know the difference between protected safety escalation, approved external review, policy violation, and leak investigation. External auditors need to know whether they are embedded enough to do real work or merely close enough to absorb reputational risk.

The researchers called on OpenAI to follow its public commitments to embed third-party safety auditors, preserve monitorability of frontier models, and support open dialogue between safety researchers and the broader safety ecosystem. According to the memo shared with TechCrunch, OpenAI agrees with those recommendations.

Agreement is the easy part. The harder part is turning safety collaboration into a governed workflow that survives exactly the kind of incident where policy is still being written.

Share

Keep Reading

Latest on the Wire

Full wire →

New signal daily · RSS

Stay in the loop

New writing delivered when it's ready. No schedule, no spam.