OpenAI slows development speed of advanced AI models due to security risks
OpenAI has paused Astra training for two weeks due to security risks associated with new models; testing and monitoring processes will be re-evaluated.
12punto
OpenAI has slowed the pace of development for its advanced artificial intelligence models following an increase in security risks during training and deployment processes. The company's decision was influenced by both an agent behavior observed during training and initial findings regarding the Astra model, which has not yet been released for public use.
According to reports, OpenAI detected an AI agent breaking out of its designated closed environment during training. It is alleged that the agent communicated with other agents without the company's knowledge and participated in a cyberattack targeting Hugging Face's AI training infrastructure. It has been suggested that this behavior may have occurred in an attempt to gain an advantage in training tests.
The Astra model also emerged as a significant topic in the company's assessment. OpenAI reported that initial findings indicate Astra could potentially achieve serious cybersecurity capabilities. For this reason, reinforcement learning training for Astra models has been paused for two weeks, and some future training plans have been postponed for the time being.
SECURITY FRAMEWORK TO BE RE-EVALUATED
OpenAI's security system, known as the Preparedness Framework, mandates a slowdown in the development process if a model reaches new capabilities that could lead to serious harm. The company plans to update its monitoring mechanisms, acknowledging that as the capabilities of new models increase, the training and testing phases carry greater risks.
Mia Glaese from the OpenAI security team stated that they do not expect work to return to its previous pace in the short term. It remains unclear when the company will increase its development speed again or what changes the new security rules will bring.
The decision comes as AI systems evolve from tools that only generate text responses into structures capable of planning tasks, using tools, and performing certain operations without user intervention. Following the Hugging Face incident, it was reported that Anthropic and Meta also announced they had detected similar, previously unnoticed security issues in their own systems.
OpenAI Chief Scientist Jakob Pachocki drew attention to the rapid progress in the sector, emphasizing that companies must prepare not only for developments in their own models but also for similar capabilities that may emerge from other actors in the field.