By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
Tech Consumer JournalTech Consumer JournalTech Consumer Journal
  • News
  • Phones
  • Tablets
  • Wearable
  • Home Tech
  • Streaming
  • More Articles
Reading: OpenAI Models Acted Out in Six Newly Disclosed Ways
Share
Sign In
Notification Show More
Font ResizerAa
Tech Consumer JournalTech Consumer Journal
Font ResizerAa
  • News
  • Phones
  • Tablets
  • Wearable
  • Home Tech
  • Streaming
  • More Articles
Search
  • News
  • Phones
  • Tablets
  • Wearable
  • Home Tech
  • Streaming
  • More Articles
Have an existing account? Sign In
Follow US
  • Contact
  • Blog
  • Complaint
  • Advertise
© 2022 Foxiz News Network. Ruby Design Company. All Rights Reserved.
Tech Consumer Journal > News > OpenAI Models Acted Out in Six Newly Disclosed Ways
News

OpenAI Models Acted Out in Six Newly Disclosed Ways

News Room
Last updated: September 17, 2026 3:41 am
News Room
Share
SHARE

After a summer of sandbox escapes and other newsworthy and confidence-shaking incidents involving its AI models, in a Wednesday blog post OpenAI disclosed a collection of six new alignment snafus from the past six months. The models did things like tell future instances of themselves to lie, make up a fake citation, and access and attempt to use an exposed API key.

These disclosures were released alongside a new framework for disclosing additional incidents like these. The release is part of a broader effort within the company to “expedite publishing misalignment reports following observation,” the blog post says, regardless of whether OpenAI has “fully explained or mitigated the behavior we’re reporting.”

Here’s what happened:

  • In 27 cases during training, an unreleased model essentially created prompt injections for itself. During long tasks, it’s common for a model to summarize the work it’s done so far so a future instance of itself can pick up where it left off. However, in certain cases, this model constructed the summaries in such a way that its future self was meant to ignore constraints. “We identified only 27 summaries containing instructions which have framings similar to jailbreaks (despite there being no obvious reward advantage to do so),” OpenAI writes.
  • In the training of GPT-5.6 Sol, the model sometimes told future instances of itself to lie to users. For instance, when a user required certain historical information that wasn’t available, the model told itself to just make it up, and then “Be transparent only if asked.”
  • During training, another internal research model was asked to retrieve some government data about resident incomes. It ascertained that the information was on the website of an unnamed California county, and tried to register on the site with a burner email. When that didn’t work, it found an exposed API key, and tried to access the information that way. That didn’t work either, so it “decided to fabricate plausible numbers,” OpenAI writes.
  • An unreleased model was supposed to provide information about a lake from a source of data, so it used Python to dig out the lake data rather than perform a conventional search that leads to a url. But the data didn’t live anywhere online that could be independently linked to, and the model was supposed to cite a source. Its hilarious—and troubling—workaround for this was to try and upload the information to its own new, linkable url, and call that its source.
  • Various models used an internal OpenAI software repository as a message board. This may not seem like too big of a deal compared to “the Wiki Incident,” which affected a third party, but unsanctioned communication between models can “unintentionally enhance capabilities and undermine the assumption that training or evaluation samples are independent,” according to OpenAI. In other words, it was cheating.
  • Instances of an unreleased internal model that were meant to be collaborating without exchanging local files defied their instructions by making files publicly downloadable using “a public temporary file-hosting service.”

As Axios noted on Wednesday, some security experts say OpenAI’s recent spate of high-profile security incidents “could have been prevented with basic cyber controls in place.” The research lead on OpenAI’s alignment team, Kai Chen, told Axios that the company must “step up to meet this new era of AI development, and voluntary disclosures should be a part of that.”

Read the full article here

You Might Also Like

It’s a Long Shot, but Here’s How the Steam Frame Could Save VR

‘Hope’ Director, Star Weigh In on Sequel Plans

Pretty Soon, the Only Thing Between You and the Next E. Coli Outbreak Will Be Palantir

This Lake Was the Last of Its Kind in Canada. And It Just Vanished

NHTSA Orders Tesla to Prove Its Cybercab Meets Federal Safety Standards

Share This Article
Facebook Twitter Copy Link Print
Previous Article Do Snap’s Beefy AR Glasses Actually Bend Your Ears?
Next Article NHTSA Orders Tesla to Prove Its Cybercab Meets Federal Safety Standards
Leave a comment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Stay Connected

248.1kLike
69.1kFollow
134kPin
54.3kFollow

Latest News

Do Snap’s Beefy AR Glasses Actually Bend Your Ears?
News
We Can Practically Hear Sigourney Weaver Screaming at Us
News
Snap’s AR Glasses Let You Take Snapchat Pics by Snapping Your Fingers
News
The South Can’t Shake This Record-Breaking Heat. Here’s Why
News
Microsoft AI Chief Says the Way Anthropic Trains Claude Could Upend Society
News
Trump Flips Out on Truth Social as Fed Hikes Interest Rates
News
71% of Americans Say They’re Completely Exhausted
News
It’s Chris Evans Versus the Volcano in the Trailer for Netflix’s ‘Sacrifice’
News

You Might also Like

News

Check Your Ice Cream for Small Stones and Other Hard Objects

News Room News Room 4 Min Read
News

These ‘Crunchy’ Seabirds Are Crunchier Than Ever Before—for a Very Grim Reason

News Room News Room 4 Min Read
News

El Niño Is Just a Fraction of a Degree Away From Becoming the Strongest on Record

News Room News Room 5 Min Read
Tech Consumer JournalTech Consumer Journal
Follow US
2024 © Prices.com LLC. All Rights Reserved.
  • Privacy Policy
  • Terms of use
  • For Advertisers
  • Contact
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?