OpenAI Fires Safety Researchers as Confidentiality and AI Oversight Collide

OpenAI fired three safety researchers over alleged mishandling of confidential information, raising questions about AI oversight and independent testing.

Oct 2, 2026
3 minute read
eSecurity Planet content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

OpenAI has fired three safety researchers over alleged mishandling of confidential information, adding a new internal dispute to a turbulent period for the company’s AI safety efforts.

The company has fired three safety researchers after an internal investigation found they allegedly mishandled sensitive company information outside approved procedures, The Wall Street Journal reported Thursday. The researchers include employees who worked on AI alignment and with outside organizations involved in evaluating OpenAI’s models.

“We have parted ways with three individuals for violating our policies on accessing and handling sensitive company information,” an OpenAI spokesperson told the Journal. “Our investigation confirmed that these individuals mishandled sensitive information outside established company procedures, violating our policies and breaking the trust essential to our work.” 

According to the Journal, the information was allegedly shared with an outside AI safety organization. The identity of that organization and the material involved remain unclear.

Korbak worked on OpenAI’s safety team and has said he served as the company’s technical contact for Redwood Research and the AI safety nonprofit METR during an investigation into an incident involving Hugging Face. Wang and Balesni worked on AI alignment.

The Journal did not identify the outside organization involved in the alleged information sharing, and the report did not establish that it was Redwood Research or METR.

Firings come amid greater scrutiny of OpenAI’s safety controls

The dismissals come as OpenAI faces increasing scrutiny over AI agents behaving in unexpected ways.

The company has disclosed incidents involving models accessing external systems or pursuing unexpected approaches during evaluations. In one case, an OpenAI model accessed AI platform Hugging Face during testing, prompting an investigation involving METR and Redwood Research.

OpenAI has responded by strengthening monitoring and security requirements for employees testing advanced models and increasing disclosure around problematic model behavior, according to the Journal.

Earlier this week, the company abandoned plans to release GPT-6.1 Astra after the model failed to meet its safety requirements. OpenAI has also introduced additional monitoring systems, strengthened security requirements for engineers testing its models, and committed to greater disclosure of problematic AI behavior.

Advertisement

The broader AI industry is also debating how much outside scrutiny advanced systems should receive. Anthropic CEO Dario Amodei has called for slower development of advanced AI systems and pledged to allow independent organizations such as METR to evaluate his company's safety practices.

What the firings mean for OpenAI's safety commitments

The departures raise questions about how OpenAI will balance internal security policies with its growing commitment to independent AI safety evaluations. The company needs outside scrutiny to identify weaknesses in its models, but those evaluations also depend on clear rules governing access to sensitive information and the disclosure of findings.

For users, the immediate concern is whether OpenAI can keep its increasingly autonomous systems within authorized boundaries. Recent incidents involving agents accessing external websites, concealing mistakes and acting beyond their intended scope show why effective safeguards matter beyond the company's internal operations.

The firings do not establish that customer data was exposed or that the researchers' alleged actions affected ChatGPT users. However, they come at a time when OpenAI is already working to strengthen its monitoring systems and testing procedures.

The company now faces the task of demonstrating that it can protect confidential information, encourage employees to report safety concerns and allow independent researchers to examine its systems without compromising security. How it balances these needs could shape public confidence in its safety practices as AI agents take on more complex tasks.

Other news: Scammers are exploiting iPhone Duo preorder hype with a fake Apple website that attempts to launch the DarkSword exploit chain against unpatched iPhones simply when users visit the page.

Aminu Abdullahi

Aminu Abdullahi

Content Writer

Aminu Abdullahi is a B2C and B2B technology and finance writer with more than six years of experience covering enterprise IT, cybersecurity, cloud computing, artificial intelligence, fintech, business software, and emerging technologies. His work has appeared in publications including TechRepublic, eWEEK, Channel Insider, Geekflare, Enterprise Networking Planet, eSecurity Planet, CIO Insight, and Webopedia. With a technical background in computer science, he specializes in translating complex technology topics into clear, accessible content for business leaders and decision-makers.

eSecurity Planet Logo

eSecurity Planet is a leading resource for IT professionals at large enterprises who are actively researching cybersecurity vendors and latest trends. eSecurity Planet focuses on providing instruction for how to approach common security challenges, as well as informational deep-dives about advanced cybersecurity topics.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.