AI Pretraining Poisoning Exposes a Shocking New Gap Safety Can’t Close
Pretraining poisoning survived fine-tuning and RLHF in a July 16 study, exposing a safety gap that post-training checks may not close.
Reporting on AI safety, covering model failures, rogue agent incidents, security testing and the rules being written to contain them.
Pretraining poisoning survived fine-tuning and RLHF in a July 16 study, exposing a safety gap that post-training checks may not close.