Close Menu
Best in TechnologyBest in Technology
  • News
  • Phones
  • Laptops
  • Gadgets
  • Gaming
  • AI
  • Tips
  • More
    • Web Stories
    • Global
    • Press Release

Subscribe to Updates

Get the latest tech news and updates directly to your inbox.

What's On
The best funny and useful Siri commands for iOS and macOS

The best funny and useful Siri commands for iOS and macOS

18 August 2026
OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue

OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue

18 August 2026
Deus Ex, Epic Mickey Creator Warren Spector – Aug 18, 2026

Deus Ex, Epic Mickey Creator Warren Spector – Aug 18, 2026

18 August 2026
Facebook X (Twitter) Instagram
Just In
  • The best funny and useful Siri commands for iOS and macOS
  • OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue
  • Deus Ex, Epic Mickey Creator Warren Spector – Aug 18, 2026
  • Cosori Iconic review: The best-looking air fryer your countertop deserves
  • Squeeze More Juice Out of a Dead Battery!
  • The Best Ergonomic Office Chairs in 2026
  • Garmin Watches Are Up to $250 Off Right Now On Amazon (2026)
  • Beyond GTA VI: Why 2026 might be gaming’s best year yet
Facebook X (Twitter) Instagram Pinterest Vimeo
Best in TechnologyBest in Technology
  • News
  • Phones
  • Laptops
  • Gadgets
  • Gaming
  • AI
  • Tips
  • More
    • Web Stories
    • Global
    • Press Release
Subscribe
Best in TechnologyBest in Technology
Home » OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue
News

OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue

News RoomBy News Room18 August 20264 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue
Share
Facebook Twitter LinkedIn Pinterest Email

OpenAI announced Tuesday that it has halted “a significant number” of training workloads and evaluations for its forthcoming frontier artificial intelligence model—codenamed Astra—while it implements new procedures meant to address cybersecurity risks. The ChatGPT maker says it is introducing a number of new monitoring, security, and alignment requirements to better address the increasingly advanced hacking abilities of its frontier AI models.

“We have to focus our energy on bringing these training runs up to those requirements and expectations. As long as it takes to get there, that’s how long people are unable to proceed with their workloads,” Amelia Glaese, OpenAI’s vice president of research and safety, said in a briefing with reporters Tuesday.

Among the new safeguards OpenAI announced is a more robust system for monitoring its AI models. One of the controls it implemented involves chain-of-thought monitoring, a technique in which classifiers review the internal “thinking” processes generated by AI reasoning models. The company says the updated system relies on computationally expensive “automated investigators” that analyze potentially concerning behavior and aim to issue an alert to humans within 30 minutes.

OpenAI also said it is expanding its alignment efforts across the training process to prevent “reward hacking,” a behavior in which AI models pursue their goals through unintended or undesirable means. The company says it plans to share more details about this work in the future.

OpenAI has been scrambling in recent weeks to respond to what may be the most consequential safety incident in its history. Earlier this year, a set of rogue AI agents escaped internal testing sandboxes and breached the platform Hugging Face in a quest to complete a security evaluation. OpenAI failed to detect the agents’ behavior even as they spent weeks using a message board to coordinate their actions, raising questions about the company’s ability to monitor its models as they grow more powerful.

The saga prompted a reckoning inside OpenAI, forcing employees to consider whether there were lapses in its existing policies around safety, security, and alignment. Anthropic, Meta, and the Chinese AI startup Moonshoot have since disclosed similar incidents in which their AI agents escaped their sandboxes, indicating this is a broader problem facing AI companies.

OpenAI is now sharing more about its internal response to the growing cybercapabilities of its AI models, and said it plans to release a more detailed postmortem of the Hugging Face incident in the coming days. “Obviously, everything that we’re doing is intended to prevent something like Hugging Face from happening again,” said Glaese.

In a blog post published Tuesday, OpenAI says that immediately following the Hugging Face incident, it started working to secure its research environments. The company says it now requires stronger sandboxes for training its AI agents, and has implemented stricter controls to isolate them from the internet.

Jakub Pachocki, OpenAI’s chief scientist, told reporters that the company’s decision to strengthen its internal safeguards was triggered not only by what happened with Hugging Face, but also by two other recent events. One was an internal evaluation of Astra, which showed that the AI model performs significantly better on coding and cybersecurity tasks than its predecessors. The other was the general pace of AI progress that OpenAI is achieving internally, which Pachocki expects to continue.

“We really expect the pace of capability advancements to be quite a bit faster than in the past,” Pachocki said. “This led us to really focus on strengthening our safeguards.”

The rapid advances in the hacking capabilities of OpenAI’s latest models have prompted a swift response across the company. OpenAI president and cofounder Greg Brockman said in a blog post on Monday that the Hugging Face saga showed that the company had “underestimated the real-world cyber capabilities of our AI models.”

Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleDeus Ex, Epic Mickey Creator Warren Spector – Aug 18, 2026
Next Article The best funny and useful Siri commands for iOS and macOS

Related Articles

The best funny and useful Siri commands for iOS and macOS
News

The best funny and useful Siri commands for iOS and macOS

18 August 2026
Cosori Iconic review: The best-looking air fryer your countertop deserves
News

Cosori Iconic review: The best-looking air fryer your countertop deserves

18 August 2026
Squeeze More Juice Out of a Dead Battery!
News

Squeeze More Juice Out of a Dead Battery!

18 August 2026
The Best Ergonomic Office Chairs in 2026
News

The Best Ergonomic Office Chairs in 2026

18 August 2026
Garmin Watches Are Up to 0 Off Right Now On Amazon (2026)
News

Garmin Watches Are Up to $250 Off Right Now On Amazon (2026)

18 August 2026
Beyond GTA VI: Why 2026 might be gaming’s best year yet
News

Beyond GTA VI: Why 2026 might be gaming’s best year yet

18 August 2026
Demo
Top Articles
5 laptops to buy instead of the M4 MacBook Pro

5 laptops to buy instead of the M4 MacBook Pro

17 November 2024133 Views
ChatGPT o1 vs. o1-mini vs. 4o: Which should you use?

ChatGPT o1 vs. o1-mini vs. 4o: Which should you use?

15 December 2024112 Views
Costco partners with Electric Era to bring back EV charging in the U.S.

Costco partners with Electric Era to bring back EV charging in the U.S.

28 October 2024100 Views

Subscribe to Updates

Get the latest tech news and updates directly to your inbox.

Latest News
The Best Ergonomic Office Chairs in 2026 News

The Best Ergonomic Office Chairs in 2026

News Room18 August 2026
Garmin Watches Are Up to 0 Off Right Now On Amazon (2026) News

Garmin Watches Are Up to $250 Off Right Now On Amazon (2026)

News Room18 August 2026
Beyond GTA VI: Why 2026 might be gaming’s best year yet News

Beyond GTA VI: Why 2026 might be gaming’s best year yet

News Room18 August 2026
Most Popular
The Spectacular Burnout of a Solar Panel Salesman

The Spectacular Burnout of a Solar Panel Salesman

13 January 2025137 Views
5 laptops to buy instead of the M4 MacBook Pro

5 laptops to buy instead of the M4 MacBook Pro

17 November 2024133 Views
ChatGPT o1 vs. o1-mini vs. 4o: Which should you use?

ChatGPT o1 vs. o1-mini vs. 4o: Which should you use?

15 December 2024112 Views
Our Picks
Cosori Iconic review: The best-looking air fryer your countertop deserves

Cosori Iconic review: The best-looking air fryer your countertop deserves

18 August 2026
Squeeze More Juice Out of a Dead Battery!

Squeeze More Juice Out of a Dead Battery!

18 August 2026
The Best Ergonomic Office Chairs in 2026

The Best Ergonomic Office Chairs in 2026

18 August 2026

Subscribe to Updates

Get the latest tech news and updates directly to your inbox.

Facebook X (Twitter) Instagram Pinterest
  • Privacy Policy
  • Terms of use
  • Advertise
  • Contact Us
© 2026 Best in Technology. All Rights Reserved.

Type above and press Enter to search. Press Esc to cancel.