Stephen Casper - Non-Consensual AI Deepfakes: AI Safety's Trial by Fire
Stephen Casper (MIT) argues that non-consensual AI deepfakes are the first real test of AI safety, and that the dominant plan has failed it. The plan is familiar: make AI safe by building safe systems. OpenAI's DALL-E 2 in April 2022 was a model example, with extensive safeguards that held. Four months later Stable Diffusion was released openly, trained without content filtering, and within two days became the tool of choice for AI-generated non-consensual content. DALL-E 2's safety work wasn't wrong, it just became irrelevant. The scale is serious. The Internet Watch Foundation recorded a 26,362% rise between 2024 and 2025 in photorealistic AI videos depicting child sexual abuse. A 2025 survey found one in eight teens know someone targeted by deepfake nudes. Most of the volume runs through a few open-weight models (Stable Diffusion, Flux, Wan 2.2) and distribution platforms (Hugging Face, CivitAI) doing little to reduce friction, despite clear evidence it works: removing adult content from training data produced roughly 1,000x lower misuse rates between Stable Diffusion 1.0 and 2.0. AI safety, Casper concludes, is not a model property but an ecosystem property. Harmful capabilities proliferate, whether through open models or closed developers willing to release them. Technical safeguards matter when they feed into standards and accountability across the ecosystem, not when they stop at one company's system. Any safety agenda that doesn't account for proliferation, he argues, should be considered unserious. Note: The opinions shared in this event are those of the speaker(s) and may not represent the views of FAR.AI or their affiliated organizations.