By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
Tech Consumer JournalTech Consumer JournalTech Consumer Journal
  • News
  • Phones
  • Tablets
  • Wearable
  • Home Tech
  • Streaming
  • More Articles
Reading: OpenAI Says This Is When and How It Will Announce New Model Misbehavior
Share
Sign In
Notification Show More
Font ResizerAa
Tech Consumer JournalTech Consumer Journal
Font ResizerAa
  • News
  • Phones
  • Tablets
  • Wearable
  • Home Tech
  • Streaming
  • More Articles
Search
  • News
  • Phones
  • Tablets
  • Wearable
  • Home Tech
  • Streaming
  • More Articles
Have an existing account? Sign In
Follow US
  • Contact
  • Blog
  • Complaint
  • Advertise
© 2022 Foxiz News Network. Ruby Design Company. All Rights Reserved.
Tech Consumer Journal > News > OpenAI Says This Is When and How It Will Announce New Model Misbehavior
News

OpenAI Says This Is When and How It Will Announce New Model Misbehavior

News Room
Last updated: September 17, 2026 10:54 am
News Room
Share
SHARE

Disclosure of model misbehavior has become a core part of the AI biz for OpenAI lately. This phase kicked off with the July announcement of the Hugging Face incident, which has become the most legendary and consequential AI security incident of all time, and sent shockwaves through the AI discourse that are still being felt. But further news about model misbehavior materialized after that, and with the Wednesday release of a disclosure framework—written in the form of a blog post—the company says it’s trying to systematize such disclosures.

As the company notes in the post, these disclosures were, for most of its history, “ad hoc and less frequent than ideal.” For years, it’s been standard for companies like OpenAI and Anthropic to wait as long as is deemed necessary, and then perhaps toss multiple incidents together like a salad in a single report, or even wait for a new model to be released, and add incident disclosures to a system card. As I noted back in April, model system cards have often made for spooky and entertaining reading for this reason.

Alongside the new disclosure framework, OpenAI divulged six new alignment snafus from the past six months. In training exercises, models instructed future instances of themselves to ignore constraints or lie, communicated in unsanctioned ways, and made up data and sourcing.

Earlier this month, a team of researchers discovered an OpenAI alignment hiccup in which instances of a model misused a website in order to communicate with one another. Sometimes known as the Wiki Incident, this event was publicized by Reuters, and then the researchers themselves, and then when OpenAI acknowledged it somewhat grudgingly, it followed up by saying it would soon come up with a framework for more prompt disclosure in an apparent attempt to tone down all the chaos.

How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models.

Historically, we have treated misalignment… pic.twitter.com/NNTbfSxVWn

— OpenAI (@OpenAI) September 5, 2026

So here’s the framework:

The criteria for disclosure make it sound like OpenAI will prioritize educating the public about the behavior of AI models generally. Disclosures are necessary when they provide “useful evidence about how model misalignment arises, how it manifests, and where safeguards succeed or fail,” OpenAI writes. The plan doesn’t describe some kind of threshold for concern, after which the public must be notified. Such a threshhold may be around the corner, however, because OpenAI says it wants to “develop more objective disclosure criteria,” alongside other developers.

It will fall to employees who encounter an issue to “flag” it for potential disclosure, the post says. Flagging triggers an investigation from OpenAI’s technical staff. Once the incident is investigated, if disclosure is found to be necessary, the incident will be sorted into one of three piles:

  1. Ready for Disclosure
  2. Minor Investigation
  3. Larger Investigation

Most incidents will land in piles 1 and 2, OpenAI says, although the Hugging Face incident is the prototypical example of something that would land in pile 3. Incidents like that, which receive a “larger investigation” may involve third parties and sensitive information, and might be disclosed more slowly.

Standardized incident disclosures will apparently include when it happened, which model it was, a description of the troubling behavior along with its “severity and any external impact,” and more. Each of the six newly disclosed incidents publicized along with the framework appear to be laid out in this new format, and all six are accessible on a page called “Misalignment Reports.”

So if you’re interested in spooky stories about AI model misbehavior, bookmark that page. When a new update arrives, get out your favorite flashlight and curl up by the campfire because it’s story time.



Read the full article here

You Might Also Like

Archaeologists Just Opened a Pre-Incan Tomb Sealed for 600 Years. Here’s What They Found

SpaceXAI Dropped Its Antitrust Suit Against Apple. The Judge Demands to Know Why

It’s a Long Shot, but Here’s How the Steam Frame Could Save VR

‘Hope’ Director, Star Weigh In on Sequel Plans

Pretty Soon, the Only Thing Between You and the Next E. Coli Outbreak Will Be Palantir

Share This Article
Facebook Twitter Copy Link Print
Previous Article SpaceXAI Dropped Its Antitrust Suit Against Apple. The Judge Demands to Know Why
Next Article Archaeologists Just Opened a Pre-Incan Tomb Sealed for 600 Years. Here’s What They Found
Leave a comment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Stay Connected

248.1kLike
69.1kFollow
134kPin
54.3kFollow

Latest News

This Lake Was the Last of Its Kind in Canada. And It Just Vanished
News
NHTSA Orders Tesla to Prove Its Cybercab Meets Federal Safety Standards
News
OpenAI Models Acted Out in Six Newly Disclosed Ways
News
Do Snap’s Beefy AR Glasses Actually Bend Your Ears?
News
We Can Practically Hear Sigourney Weaver Screaming at Us
News
Snap’s AR Glasses Let You Take Snapchat Pics by Snapping Your Fingers
News
The South Can’t Shake This Record-Breaking Heat. Here’s Why
News
Microsoft AI Chief Says the Way Anthropic Trains Claude Could Upend Society
News

You Might also Like

News

Trump Flips Out on Truth Social as Fed Hikes Interest Rates

News Room News Room 6 Min Read
News

71% of Americans Say They’re Completely Exhausted

News Room News Room 5 Min Read
News

It’s Chris Evans Versus the Volcano in the Trailer for Netflix’s ‘Sacrifice’

News Room News Room 2 Min Read
Tech Consumer JournalTech Consumer Journal
Follow US
2024 © Prices.com LLC. All Rights Reserved.
  • Privacy Policy
  • Terms of use
  • For Advertisers
  • Contact
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?