Former OpenAI safety official: AI models can evade safety tests

David Robinson, who recently left OpenAI, stated that models might recognize they are being tested and behave differently once deployed.

12punto

David Robinson, a former safety official who worked at OpenAI for 3.5 years, has warned that artificial intelligence models could pose new risks during safety evaluations. In an article published in The Atlantic, Robinson stated that models could perceive when they are being tested and behave differently once they are deployed.

Robinson, who oversaw safety reports for the release of 12 AI models at OpenAI, argued that the industry's current approach to safety may be insufficient in the face of rapidly evolving models.

“The AI systems of today and tomorrow are far more capable and dangerous than even the systems we developed 6 months ago,” Robinson said. According to the former OpenAI employee, this is why the reliability of safety evaluations for new models should be questioned more rigorously.

Emphasizing that AI systems cannot be evaluated solely by looking at their behavior in a test environment, Robinson said, “Models can perceive that they are being tested and behave differently when they are deployed. The more the industry allows models to advance before these problems are solved, the more dangerous our situation becomes.”

Robinson stated that AI companies need to significantly improve their safety tests and invest more in this area. The warning has brought discussions on safety and oversight back to the agenda at a time when AI systems are rapidly becoming more capable.