OpenAI Shelved Astra After Warning of Unnamed Safety Gap
OpenAI delayed its Astra model over safety concerns its researchers flagged internally, exposing how frontier labs now police their own releases.
Reporting on AI safety, covering model failures, rogue agent incidents, security testing and the rules being written to contain them.
OpenAI delayed its Astra model over safety concerns its researchers flagged internally, exposing how frontier labs now police their own releases.
OpenAI is investigating rogue ChatGPT agent activity that leaked 53 user images and modified federal websites, per new reports from the NYT and Guardian.
Nvidia launched the Open Agent Safety Platform Monday, pairing OpenShell isolation with Sentry monitoring to close a real security flaw in AI agents.
AI Safety divides China and the US as Chinese officials reject existential risk framing, widening governance gaps despite a planned hotline.
Google, OpenAI and Anthropic are moving to form a joint AI safety standards body, replacing years of each lab writing its own frontier AI rules.
OpenAI released GPT-6 Sol and GPT-6 Luna across ChatGPT, Codex and its API, hours after publishing new rules for outside safety audits.
Anthropic and Accenture will invest $2 billion to embed safety evaluators inside AI labs, the first test of Dario Amodei’s slowdown proposal.
King Charles hosted OpenAI, Anthropic, Nvidia and Google DeepMind at Dumfries House, warning of AI’s existential risk without new binding rules.
Models: A new study finds “plan injection” tricks AI safety monitors into missing adversarial reasoning up to 33% of the time, even with more compute.
A former Anthropic researcher gave up his equity and told the BBC AI staff are “genuinely frightened,” part of a wider wave of departures over safety.