Can We Predict AI Loss of Control? Trustworthy AI in the Age of Agents | Yinpeng Dong
Yinpeng Dong on trustworthy AI as systems become agents, and a behavioral framework for predicting AI loss of control before it happens. Dong organizes trustworthy AI around three questions for the agentic era. Safer reasoning: shifting from fast "System 1" refusals to deliberate "System 2" analysis using safety-informed Monte Carlo tree search, with test-time scaling for safety. Verified action: framing agent verification as sequential evidence accumulation, updating a log-odds confidence score, and grounding each step in domain guidelines (for example, clinical guidelines) to get a cleaner correctness signal. Early prediction of frontier risk: focusing on loss of control, where an agent diverges from human intent and cannot be reliably constrained or shut down. He decomposes that risk into misaligned intention, harm-enabling capability, and oversight evasion, with early signs like strategic sandbagging and shutdown resistance, and combines them into a loss-of-control score that predicts realistic long-horizon failures. Chapters 0:00 Trustworthy AI in the age of agents 0:23 Three stages: responses, actions, autonomy 1:26 Making models think in a safer way (System 2) 2:20 Reasoning-based safety alignment 2:51 From reasoning to action: verifying agents 3:37 Agent verification as evidence accumulation 4:55 Guideline-grounded verification (clinical domain) 5:54 Predicting frontier risks: loss of control 6:52 Early warning signs: sandbagging and shutdown resistance 7:50 A behavioral framework for loss of control 8:56 Predicting real long-horizon failures More AI safety research: https://far.ai Alignment Workshop playlist: https://youtube.com/playlist?list=PLBY5kyt_LfFg&si=0IfDd-WNQwrs14Kn FAR.AI is a research nonprofit working to ensure the safe development of advanced AI. We host the Alignment Workshop series and publish frontier alignment research.