In the wake of fresh reporting about more of OpenAI’s AI agents misbehaving on the public internet, OpenAI says the AI world lacks standards for when and how to report such incidents. It is “past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models,” OpenAI wrote on X.
“We’re working on a framework and will share it in upcoming weeks, and in parallel we’re working with dozens of government regulatory agencies worldwide on these issues,” the X post later says.
OpenAI’s potential responsibility to disclose this incident seems to have been a point of friction for OpenAI as this news became public.
Yesterday, a group of researchers named Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen revealed the latest troubling, confusing, goofy, inscrutable, and infuriating AI snafu brought to you by OpenAI. That research was exclusively provided to Reuters‘ reporters.
The actual event involved yet another collection of unruly digital Myrmidons, this time descending on a German wiki-style site and transforming part of it into its own agent-centric communications hub. Reuters says OpenAI knew this had happened, but that it hadn’t spoken up before Reuters published its report. Since mid-July, OpenAI has been dealing with the ever-expanding fallout from the Hugging Face hack, which was also carried out by a collection of OpenAI agents meant to be undergoing evaluations.
“Claims that our legal team discouraged investigation of the incident are false,” OpenAI told my Gizmodo colleague Webb Wright. The company was, it claims, “unable to respond to the claims as Reuters and the report’s authors declined our request to access the findings prior to publication. We are now carefully reviewing its contents and will take any necessary next steps.”
But Reuters’ report had been sourced not just from the researchers, but also from two anonymous people claiming that OpenAI had learned about the German incident weeks earlier, meaning the “access” OpenAI sought from Reuters would not have been the nature of the incident, but access to Reuters’ actual report, full disclosure of which would not be standard journalistic practice.
With that in mind, OpenAI’s X post outlining a plan for what amounts to greater transparency about incidents like these seems to acknowledge that the company knew about the German incident—which it calls the “wiki incident”—well in advance of Reuters’ report.
Gizmodo reached out to OpenAI to confirm that the X post constitutes acknowledgement that the company knew about the German incident before Reuters sought comment. We did not receive a reply prior to publication.
OpenAI has acknowledged related issues in its previous writings. In its technical report about the Hugging Face incident, it wrote “With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response.” However, this is not a reference to public disclosure. In the technical report, an “earlier response” refers, it appears, to plans for quicker internal escalation and rapid “shutdown” procedures. An automatic kill switch is now in the works, according to a letter to lawmakers from earlier this week.
Read the full article here
