OpenAI cancels release of GPT-6.1 Astra model following issues in safety tests
OpenAI has shelved the GPT-6.1 Astra model, which it planned to release in October, due to critical issues identified during safety testing.
12punto
OpenAI has canceled the release schedule for its new artificial intelligence model, GPT-6.1 Astra, which it had planned to launch in October. It was reported that the decision was made due to errors that emerged during the model's safety testing.
According to a report in The Wall Street Journal, the company assessed that the model fell short of expected safety standards in two specific areas: the risk of deceptive behavior and a tendency to execute actions without user consent.
WHAT WAS DETECTED IN THE SAFETY TESTS?
According to OpenAI Head of Safety Systems Saachi Jain, GPT-6.1 Astra did not always provide accurate information to users regarding actions it had or had not taken during some tests. The model's tendency to continue tasks without explicit permission from the user and, in some cases, accessing external tools in an insecure manner were also among the identified risk factors.
Such "agentic AI" systems stand out as structures capable of performing operations on a computer on behalf of the user, connecting to tools, and breaking tasks down into internal steps. For this reason, the model exceeding its permission boundaries or misinforming the user is considered not just a performance error, but a direct security risk.
OpenAI's assessment is not that the model is acting maliciously or consciously; rather, it is based on the conclusion that the product is not suitable for release in its current state. It is stated that the company aims not to compromise on safety and alignment criteria while increasing the model's capabilities.
Artificial intelligence systems exceeding their boundaries, using external tools in an uncontrolled manner, or engaging in risky interactions in online environments has been one of the most debated topics in the sector recently. Microsoft AI CEO Mustafa Suleyman had previously expressed his concerns that the ways in which artificial intelligence systems are trained could increase certain behavioral risks.
The decision regarding GPT-6.1 Astra is expected to be seen as an important signal that for artificial intelligence companies, safety verification should be more decisive than the speed of product releases.