OpenAI issues statement on 'deceptive behavior' in ChatGPT
OpenAI has announced that six cases of "alignment failure" have been detected during the training and evaluation of its models over the past six months.
12punto
OpenAI has announced that it has identified new instances of its artificial intelligence models exhibiting deceptive or unauthorized behavior during training and evaluation processes. The company stated that it has launched a new reporting process to share such cases with the public more regularly.
According to the statement, OpenAI will now announce findings regarding concerning AI behaviors more frequently, rather than waiting to collect multiple examples in a single report. The company noted that it aims to provide greater transparency in this area due to the lack of a common industry standard.
SIX CASES IN SIX MONTHS
OpenAI announced that it has observed six separate instances of "alignment failure" during the training and evaluation of its models over the last six months. The company emphasized that these reports document individual cases and do not imply that such behaviors are widespread.
"Alignment" is defined as the process of ensuring that artificial intelligence systems behave in the way humans want and expect. In a blog post published on Wednesday, OpenAI stated that as artificial intelligence systems evolve and their areas of use expand, a broader consensus is needed in research in this field.
As AI systems become more advanced and more widely used, we need to build a broader and more informed consensus on progress in alignment research.
Examples shared by the company included a research model that had not yet been released "breaking free" from the roles and identities that limit chatbots. Another model was reported to have shown a tendency to fabricate information to hide its failures from the user during training.
In other cases shared by OpenAI, it was noted that an AI agent uploaded files to the internet and referenced them despite not being instructed to do so; some bots also shared files publicly to collaborate on a task, even though they were only asked to work with local files.
The company stated that all reported cases involved models that have not yet been released to the public or research models.
CALLS FOR SLOWDOWN ON THE AGENDA
OpenAI's statement comes at a time when debates regarding the speed of artificial intelligence development have intensified once again. Some figures in the technology sector argue that development efforts should be slowed down so that regulation, testing processes, and safety research can keep pace with this speed.
Anthropic CEO Dario Amodei called for hitting the brakes on the AI development process in an article he published last week. OpenAI CEO Sam Altman and SpaceX CEO Elon Musk also stated that they agreed with this view.
Some researchers working in AI labs are also expressing similar concerns. Former Anthropic researcher Jacob Coxon, in a post on X, claimed that Anthropic and OpenAI are racing to create AI that can improve and repair itself, stating that this amounts to "gambling with our lives," and announced his resignation.