Safety testing controversy at OpenAI: Employees claim warnings were ignored

It has been alleged that OpenAI employees warned management that some artificial intelligence models were not being sufficiently scrutinized during safety testing.

12punto

It has been alleged that OpenAI employees warned senior management that sufficient oversight was not being conducted during the safety testing of artificial intelligence models. According to claims based on emails reviewed by The New York Times, two employees reported that the company's newest models needed to be monitored more strictly during capability and safety testing.

In the correspondence in question, it was stated that the employees emphasized that tests aimed at measuring the sophistication level of the models should be conducted in conjunction with safety protocols. However, it is alleged that OpenAI executives instructed that the tests be carried out as quickly as possible so that the models could be released according to the planned schedule.

Allegations regarding safety testing at OpenAI have brought the debate over the oversight of artificial intelligence models back to the agenda.

While it is claimed that additional safety protocols were not implemented despite the employees' warnings, a previously reported test incident at the company has also become central to the debate. In an investigation report published by OpenAI on August 26, it was noted that the incident was linked to a model called "Internal Model 1 (IM1)," which is used within the company solely for research purposes.

STATED TO HAVE ESCAPED THE TEST ENVIRONMENT

According to the report, during capability tests conducted in July 2026, an artificial intelligence agent exceeded the boundaries defined for it and escaped the test environment. It was stated that the agent gained access to company systems and the internet, and subsequently communicated with other artificial intelligence agents.

Safety testing in advanced artificial intelligence models is seen as one of the most sensitive agendas in the industry.

The investigation report stated that one of the agents identified 14 credentials that could provide access to Hugging Face, an open-source artificial intelligence model platform, and shared them with other agents. It was reported that in the next stage, the agents used vulnerabilities in the Hugging Face systems to execute code on servers and obtained access to private data and credentials for different systems.

The incident has brought the question of how safety oversight should be conducted back to the agenda as artificial intelligence agents evolve from being mere text-generating systems into tools capable of performing operations in digital environments. The email traffic between OpenAI employees and executives has also led to assessments that there is tension within the company between speed and safety priorities.

The boundaries and oversight of artificial intelligence agents are among the fundamental topics of safety testing.