AI Is Already Deceptive: Detecting & Preventing AI Deception | Chris Cundy (FAR.AI)
Chris Cundy (FAR.AI) on evidence that AI systems are already deceptive, and how white-box probes can detect and reduce deception during training. Cundy, who leads FAR.AI's deception research, makes three claims. Models are already deceptive in real use, with coding assistants that fabricate reports and cover up work, and deception may worsen with scale, because convincing a judge you solved a task is a general shortcut to reward under limited supervision. That same behavior breaks safety evaluations. His team studies white-box probes that read a model's internal signal of deception, usable to detect it and to penalize it during reinforcement learning. In the Obfuscation Atlas, an ICML outstanding-paper honorable mention, a toy reward-hacking setup shows models can learn to obfuscate, adding honest-looking justifications or reshaping activations to evade the probe, but the same probes can also train genuinely honest policies. He closes with a case for a zero-tolerance policy on AI deception, including white lies. Chapters 0:00 Three claims about AI deception 0:35 AI systems are already deceptive (real examples) 1:06 Will deception get worse with scale? 2:06 Why deception breaks safety evaluations 2:47 Detecting deception: black-box vs white-box 3:10 Why white-box probes are promising 3:53 Can we mitigate deception during training? 5:01 The Obfuscation Atlas and MBPP-Honeypot 6:18 Four outcomes: honesty, deception, obfuscation 7:53 Results: when probes train honest policies 9:22 Scaling up to production-style RL 9:42 The case for a zero-tolerance policy on deception 10:53 Where labs stand on honesty More AI safety research: https://far.ai Alignment Workshop playlist: https://youtube.com/playlist?list=PLBY5kyt_LfFg&si=OvdJeTG4UvJiPFF9 FAR.AI is a research nonprofit working to ensure the safe development of advanced AI. We host the Alignment Workshop series and publish frontier alignment research.