Is Your AI Model Actually Secure? The Jailbreak Problem | Adam Gleave (FAR.AI)
Adam Gleave (FAR.AI) on why some frontier models resist every jailbreak tested while others break in hundreds of ways across high-risk domains. FAR.AI tested the safeguards of leading frontier models against 1,500 jailbreak combinations. OpenAI's and Anthropic's latest models resisted nearly every attack, while Grok and Gemini produced hundreds of universal jailbreaks that unlocked chemical, nuclear, radiological, explosives, and cyber capabilities. Every model held in biosecurity. The talk maps these jailbreaks to misuse already documented by the UK AI Security Institute, Google, and Cambridge University researchers, and explains why specialized classifiers make stronger safeguards cheap and achievable for any developer. Chapters 0:00 Uneven safeguards across frontier models 0:22 The upside: what capable models deliver 0:43 The dark side: AI misuse happening now 2:19 Testing safeguards: 1,500 jailbreak combinations 2:52 Results by domain: what held and what broke 3:33 The fix: cheap, documented safeguards 4:20 A minimal standard, and what comes next More AI safety research: https://far.ai FAR.AI is a research nonprofit working to ensure the safe development of advanced AI. We host the Alignment Workshop series and publish frontier alignment research.