Making Open-Weight AI Safer by Filtering Training Data | Stella Biderman (EleutherAI)
Stella Biderman (EleutherAI) on a safety method built for open-weight models: filtering dangerous knowledge out of the pretraining data. Biderman argues that the standard safety toolkit, input and output filtering, alignment fine-tuning, and know-your-customer checks, was designed for API models and does not hold for open-weight models that anyone can fine-tune. The alternative, from her team with Oxford and the UK AISI: because narrowly useful dangerous knowledge is intellectually isolated, it can be removed from the pretraining data. Training from scratch with this filtering sharply lowers biorisk capability, measured on WMDP-Bio, and stays robust under adversarial fine-tuning, roughly an order of magnitude more robust than other methods. She gives three conditions for when filtering works and expects it to help most for CBRN information. Chapters 0:00 Making open models safer through data filtering 0:17 Why standard safety methods fail for open-weight models 1:12 Designing safety around open-model developers' needs 1:39 The idea: filter dangerous knowledge from training data 1:53 Results: a big drop in biorisk capability (WMDP-Bio) 2:23 Robustness to adversarial fine-tuning 3:03 The cost vs. robustness trade-off 3:38 When does data filtering actually work? Three conditions 4:44 Which domains fit, and CBRN 5:18 Scaling and low-hanging fruit More AI safety research: https://far.ai Alignment Workshop playlist: https://youtube.com/playlist?list=PLBY5kyt_LfFg&si=B9I56daxBmDAeRwQ FAR.AI is a research nonprofit working to ensure the safe development of advanced AI. We host the Alignment Workshop series and publish frontier alignment research.