Mitigating Power Concentration and Multi-Agent Risks in AI | Samuel Simko (EuroSafeAI)
Samuel Simko (EuroSafeAI) on why AI safety has to extend beyond single models, to multi-agent interactions and risks to democratic institutions. Simko presents EuroSafeAI's work across three areas. Model-level safety: making models robust to jailbreaks and misalignment, including a defense that uses "honeypot" replies to make attacks rarer and less useful, plus stronger automated safety judges. Multi-agent safety: game-theoretic modeling (GT-HarmBench) showing that models which look aligned in isolation can produce serious risks when they interact, sometimes escalating in arms-race dynamics. Societal risk: power concentration and threats to democracy, with benchmarks like SocialHarmBench probing sociopolitical attacks. His argument is that research clusters on the first two areas and gives the third far less attention than it needs. Chapters 0:00 Intro: EuroSafeAI and a European perspective 0:37 Three pillars of AI safety 0:50 Model-level safety and honeypot defenses 1:35 Building better safety judges (JudgeStressTest) 2:02 Multi-agent safety (GT-HarmBench) 2:40 When aligned models escalate: arms-race dynamics 2:59 Risks to democracy and power concentration 3:26 SocialHarmBench: sociopolitical attacks 3:56 Position paper: AI risks to democratic systems 4:26 Takeaway: safety needs all three dimensions More AI safety research: https://far.ai Alignment Workshop playlist: https://youtube.com/playlist?list=PLBY5kyt_LfFg&si=B9I56daxBmDAeRwQ FAR.AI is a research nonprofit working to ensure the safe development of advanced AI. We host the Alignment Workshop series and publish frontier alignment research.