By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
Tech Consumer JournalTech Consumer JournalTech Consumer Journal
  • News
  • Phones
  • Tablets
  • Wearable
  • Home Tech
  • Streaming
  • More Articles
Reading: OpenAI Says Humans Need to Be Able to Monitor How AI ‘Thinks.’ Astra Makes That Much Harder
Share
Sign In
Notification Show More
Font ResizerAa
Tech Consumer JournalTech Consumer Journal
Font ResizerAa
  • News
  • Phones
  • Tablets
  • Wearable
  • Home Tech
  • Streaming
  • More Articles
Search
  • News
  • Phones
  • Tablets
  • Wearable
  • Home Tech
  • Streaming
  • More Articles
Have an existing account? Sign In
Follow US
  • Contact
  • Blog
  • Complaint
  • Advertise
© 2022 Foxiz News Network. Ruby Design Company. All Rights Reserved.
Tech Consumer Journal > News > OpenAI Says Humans Need to Be Able to Monitor How AI ‘Thinks.’ Astra Makes That Much Harder
News

OpenAI Says Humans Need to Be Able to Monitor How AI ‘Thinks.’ Astra Makes That Much Harder

News Room
Last updated: September 5, 2026 2:57 am
News Room
Share
SHARE

OpenAI released GPT-6 Astra on Thursday, describing it as “the world’s most intelligent and aligned model.” Company president Greg Brockman went further, saying it would be remembered as the world’s first genuine glimpse of artificial general intelligence—the dawn of a brave new world where computers are more intelligent than humans. 

What he didn’t mention is that with a jump in intelligence comes a greater difficulty in understanding how those systems work. That could be a serious problem moving forward, as AI agents continue to go rogue and government guardrails are nowhere in sight.

The past couple of weeks have been particularly dramatic for OpenAI. Which is really saying something, considering the company’s entire lifespan has been one controversy after another. On Tuesday, less than a week after independent firms Redwood Research and METR published their investigations into the recent Hugging Face hack, The Information reported that Astra had been partially developed using a technique that can make AI more capable, but also obscure its reasoning process—the steps it takes to solve a particular problem, including any dangerous missteps it makes along the way. 

Those steps are traditionally recorded in chain-of-thought (CoT) transcripts, which are basically the model’s complex pattern-detection process translated from an opaque machine language into plain English, or at least something close. It’s far from perfect, but it’s at least a rough window into how an AI model “thinks.” It was also essential to the third-party researchers who uncovered how OpenAI’s agents were able to secretly mass into a “swarm” and breach Hugging Face. The lack of CoT transcripts “would have greatly undermined our investigation,” Ryan Greenblatt, the chief scientist at Redwood Research and the leader of the nonprofit’s probe into the Hugging Face hack, wrote in an X post on Tuesday.  

Many people were alarmed, therefore, by The Information’s report that OpenAI was moving ahead with a technique that would make their AI systems even more of a black box. “This may be the single worst development for AI security/safety to date,” Greenblatt said in his X post. 

OpenAI’s own chief scientist, Jakub Pachocki, said the reporting had been “confused,” but he didn’t get more specific or deny the company’s use of recurrent depth to train Astra. In the middle of last year, Pachocki and multiple other OpenAI researchers were listed as coauthors on a paper which argued CoT was a valuable but “fragile” mechanism for keeping an eye on the behavior of AI agents. OpenAI has also said that monitoring CoT has helped its own researchers cut back on models’ misaligned behavior. And in a blog post published the same day as The Information’s report, the company said, vaguely, that its new model would be deployed “with additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions.”

The model’s safety card is not reassuring on that front. According to the company’s own internal tests, “GPT-6 Astra shows a substantial decrease in chain-of-thought monitorability compared to previous models.” Tests also found that Astra was more likely than its predecessors to change its note-taking process when it knew it was being graded: “In one of our evaluations, we find that simply informing the model that a CoT monitor is present can result in the model disproportionately shortening its CoT,” OpenAI wrote in the system card.

But all this is secondary, according to the company, since its internal tests also showed that Astra was less likely to try to evade the cybersecurity restrictions placed upon it, “which make us confident in still deploying this model to the wider public.” The system card added that OpenAI “will not accept further degradation of monitoring beyond a limit,” without elaborating on how such a limit might be defined.

OpenAI alignment researcher Tomek Korbak has said that the decrease in monitorability was a byproduct of the models themselves becoming more intelligent, rather than due to “architecture changes”—almost certainly a reference to recurrent depth. Later in the same thread, he said he was “deeply worried” by the prospect of losing CoT as models evolve. “CoT monitoring is a core part of our misalignment safety strategy that has no good substitute now,” he wrote.

Read the full article here

You Might Also Like

‘Silo’ Season 4 Preview Teases Major Confrontation and Reveals 2027 Release Date

Hasbro Is Looking for Another Studio to Make ‘Dungeons & Dragons 2’

Zohran Mamdani Is So Popular He Offers Influencers Something More Valuable Than Money: Clout

‘Hope’ Director Na Hong-jin Wanted to Make an Alien Movie From a Cosmic Perspective

UN Approves Resolution for World Map That Depicts Africa’s Size More Accurately

Share This Article
Facebook Twitter Copy Link Print
Previous Article Hasbro Is Looking for Another Studio to Make ‘Dungeons & Dragons 2’
Next Article ‘Silo’ Season 4 Preview Teases Major Confrontation and Reveals 2027 Release Date
Leave a comment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Stay Connected

248.1kLike
69.1kFollow
134kPin
54.3kFollow

Latest News

Zach Cregger Offers a Few Hints About His Sci-Fi Film ‘The Flood’
News
Epinephrine Allergy Injections Recalled Over ‘Particulate Matter,’ Including Glass
News
Pete Hegseth’s Testosterone Push for US Troops Is Off to a Very Low-T Start
News
This ‘Super’ El Niño Has Whipped Up a Never-Before-Seen Cyclone Trio in the Pacific
News
CNBC Accidentally Admits Why Employers Love AI
News
George R.R. Martin’s Review of the ‘Game of Thrones’ Play: ‘A Splendid Time’
News
Scientists Found a Massive Ice Reservoir Hidden Beneath Utah’s Rocky Mountains
News
Extreme Drought Has Shrivelled New Mexico’s Largest Reservoir Down to Just 1.4%
News

You Might also Like

News

This Startup Wants to Reach Alpha Centauri in 80,000 Years. And It’ll Only Cost $15 Million

News Room News Room 5 Min Read
News

Would You Buy an AI Smart Home Hub for $20,000?

News Room News Room 3 Min Read
News

Reolink Debuts OMVI 2i Ultra, Home Hub 2, and TrackFlex WiFi at IFA 2026, Expanding Key Product Lines and Driving Innovation

News Room News Room 9 Min Read
Tech Consumer JournalTech Consumer Journal
Follow US
2024 © Prices.com LLC. All Rights Reserved.
  • Privacy Policy
  • Terms of use
  • For Advertisers
  • Contact
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?