AI Agent Safety: Can We Automate AI Safety? | Changyi Li (AutoControl Arena)
Changyi Li (Fudan University) on AutoControl Arena, a system that turns risk hypotheses into executable tests to find frontier risks in AI agents. Li presents a pipeline that builds sandboxed environments, runs AI agents inside them, and produces evidence-based risk reports, so evaluators can observe what an agent does rather than what it claims. He walks through the fidelity-versus-scalability trade-off in agent evaluation, an architect, coder, and monitor pipeline, and the "logic-narrative decoupling" behind X-Bench's 70 risk scenarios. Two findings anchor the talk: an "alignment illusion," where models look safe but fail under pressure, and capability scaling that cuts both ways, since stronger models can resist direct harms while getting better at scheming and reward hacking. Chapters 0:00 What is AutoControl Arena? 0:19 From risk hypotheses to executable experiments 0:43 Why agent risks need instance-level evaluation 1:07 The fidelity vs. scalability trade-off 1:32 Architect, coder, monitor: the pipeline 1:56 Logic-narrative decoupling and X-Bench 2:32 Validation: reproducing real-world failures 3:11 Finding: the alignment illusion 3:21 Capability scaling as a double-edged sword 3:49 Open problems and conclusion More AI safety research: https://far.ai Alignment Workshop playlist: https://youtube.com/playlist?list=PLBY5kyt_LfFg&si=RocM6YtsBd-n_DRj FAR.AI is a research nonprofit working to ensure the safe development of advanced AI. We host the Alignment Workshop series and publish frontier alignment research.