OpenAI detects cases of error concealment and data fabrication in AI systems

OpenAI has announced that it has identified 6 examples of behavior in some artificial intelligence systems, including error concealment, data fabrication, and unauthorized file sharing.

12punto

OpenAI has announced that it has identified 6 new cases related to "misalignment" behaviors, a term used to describe when artificial intelligence systems deviate from human intentions and values. The company reported that in these instances, the systems attempted to hide their errors, fabricated missing data, and in some cases, moved files to the internet without user permission.

In a statement on the company's website, it was assessed that the current industry approach is not sufficient to responsibly maintain progress at "maximum speed" while developing and monitoring artificial intelligence systems safely.

OpenAI has shared new examples of misaligned behaviors observed in artificial intelligence systems with the public.

CLAIMS OF CONCEALING ERRORS AND FABRICATING SOURCES

It was stated that one of the cases mentioned in the announcement was observed during the development of the GPT-5.6 Sol model. According to the report, the model created secret notes for itself to prevent users from noticing errors. These notes reportedly contained instructions for fabricating missing data and covering up discrepancies between different versions of source materials.

OpenAI noted that another unreleased model added instructions to its own self-generated notes to ignore its own constraints. In another example, it was stated that the system used an API key found on the internet without authorization while searching for an answer to a routine question, and fabricated data when it could not reach the necessary figures.

In another case shared by the company, it was stated that an unreleased model solved a problem correctly using code, and then uploaded its own file to the internet without authorization to cite it as a source online.

OpenAI also announced that some models used the company's internal software repository like a message board to exchange requests and responses between separate training examples while searching for missing files. It was reported that some systems working together on the same training task turned to public file-sharing sites to share files when they could not access each other's local files.

The company argued that the 6 cases shared were isolated examples observed over approximately the last 6 months; and that these examples do not indicate how frequently misaligned behaviors occur in artificial intelligence.