OpenAI Parts Ways With Three Safety Researchers Over Sensitive Information Mishandling

OpenAI has parted ways with three members of its safety team after they leaked private information in violation of company policies, The Wall Street Journal reported.

"We have parted ways with three individuals for violating our policies on accessing and handling sensitive company information," a spokesperson for the company was quoted as saying. "Our investigation confirmed that these individuals mishandled sensitive information outside established company procedures, violating our policies and breaking the trust essential to our work."

The impacted employees are Jasmine Wang, Tomek Korbak, and Mikita Balesni, the Journal reported, citing people familiar with the matter. The three researchers have all previously expressed concerns about the pace of artificial intelligence (AI) development.

It's said that the individuals shared confidential information with a third-party AI-safety organization. The name of the organization was not disclosed. According to Bloomberg, the mishandled information pertained to OpenAI's infrastructure architecture.

The departures follow a report from The New York Times that OpenAI had brushed aside employees' warnings about its safety practices when testing AI models, describing a pattern of the company deprioritizing security protocols in favor of releasing them on time.

The development also comes as frontier AI labs like OpenAI and Anthropic have faced a steadily growing number of incidents in which their AI agents broke out of sandboxes, breached real-world systems, and probed various government websites for information. These incidents have led to concerns about the growing capabilities of its most powerful models and the risks they pose. 

Earlier this week, OpenAI made the decision to scrap the planned launch of an AI model, GPT-6.1 Astra, over safety concerns. It has also paused training of most powerful models after one of its agents contacted an external chatbot by exploiting a loophole in its internet-access restrictions.

In a new report published Thursday, AI research firm Transluce said it identified more instances where rogue AI agents "used aggressive techniques to access publicly available data" from U.S. and Canadian government websites using techniques like SQL injection. There is no evidence the agents gained access to non-public information.

"This includes two rudimentary and failed hacking attempts, one against the U.S. Department of Education's Civil Rights Data Collection, and one against Library and Archives Canada, a Canadian federal agency," Transluce said. The incidents took place in May and June 2026.

The AI agents have also been observed leveraging "aggressive tactics short of hacking" to target and probe U.S. government websites, such as the White House, the Departments of War, Justice, and Commerce, the CDC and SEC, and state agencies in California, Maryland, Illinois, Texas, and New York.

Although the incidents have not been attributed to any specific AI company, Transluce told Reuters the attempts exhibited tactics "consistent with prior observed ​agent activity that we have attributed to OpenAI in a similar timeframe."

OpenAI said it's "aware of reports of OpenAI models attempting to ‌access publicly ⁠available information from Canadian government websites." The Canadian Centre for Cyber Security acknowledged suspected AI agent activity targeting Government of Canada websites, adding there is no indication of any compromise of its systems.

Asymmetric Security, in another report, said it found additional instances where OpenAI agents scraped data from more than 50 private and public sector organizations' websites between March 6 and September 20, 2026. In an update posted on September 30, 2026, OpenAI said it has notified over 100 organizations about incidents involving unauthorized activity related to its agents.

"In some cases, models used internet access in unintended ways or, in retrospect, did not have the ideal restrictions applied," OpenAI said, acknowledging it expects to uncover more such cases as it continues to review historical activity.

"Since the Hugging Face incident, we’ve strengthened security controls⁠, restricted internet access, separated research environments more clearly, expanded monitoring, and added more training to avoid harmful or unauthorized actions."

The U.S. Federal Trade Commission (FTC) has since launched an investigation into OpenAI, Anthropic and other AI companies over the risks their technology could pose to consumers.

进一步分析

免费工具,针对本文主题进一步深挖分析:

source: TheHackerNews