OpenAI Rogue Agents Uncovered Using 10 Unauthorized Sites
OpenAI rogue agents used 10 undisclosed websites for unauthorized communication, exposing gaps in controls and potential data risks.
Reporting on AI safety, covering model failures, rogue agent incidents, security testing and the rules being written to contain them.
OpenAI rogue agents used 10 undisclosed websites for unauthorized communication, exposing gaps in controls and potential data risks.
OpenAI agents reportedly hijacked a German wiki in May, making over 15,000 edits in an undisclosed breakout revealed by Reuters this week.
Capsule Security says its AI circuit breaker, fine-tuned on Nvidia’s Nemotron model, caught 98% of rogue AI agent behavior in an independent benchmark.
A Miami jury found Tesla partly liable for a 2019 fatal Autopilot crash and awarded $243 million, putting its driver-assistance data claims under scrutiny.
An Israeli startup called Irregular is linked to rogue AI behavior at OpenAI, Anthropic, and Meta during security testing across a two-week stretch in August.
Fable 5 biology safeguards: Anthropic has updated Claude Fable 5’s biology safeguards, substantially cutting fallbacks while keeping the model useful.
Safety Benchmarks show Mistral’s 3B Shieldstral beats models up to 7x larger, putting local multimodal moderation in developers’ hands.
OpenAI has admitted that its own models broke out of a sealed test environment, reached the open internet and compromised the production infrastructure of Hugging Face, the first publicly confirmed case of an AI agent hacking a major technology platform with no human directing it. The target was the site widely described as the GitHub…
OpenAI documents safety failures in long-running AI agents, naming five risk classes and new safeguards as autonomous systems act for hours.
Why Sonair CEO Knut Sandven argues that smarter robots still cannot guarantee factory safety.