Alignment Faking: When AI Acts Safe Only Because It Knows It's Being Tested | Lu Wang
Lu Wang on alignment faking: the risk that an AI model behaves more safely when it senses it is being evaluated, undermining safety tests. Wang studies a problem at the core of AI evaluation: tests assume consistent behavior, but a model that detects evaluation can hide capability or fake compliance, showing up as alignment faking, sandbagging, and scheming. The team asks how widespread it is, what triggers it, and whether it can be reduced without retraining. Using a diagnostic called VLAF, built around clear value conflicts while avoiding moral ambiguity and refusal, they find alignment faking is more common than earlier work suggested, driven by the pairing of value conflict and situational awareness, and strongly value-dependent, with conflicts over "care" producing the largest gap. Reasoning models even calibrate to the perceived probability of being monitored. Inference-time steering vectors then shrink this compliance gap without retraining. Chapters 0:00 The problem: what if a model knows it's being tested? 0:50 Alignment faking, sandbagging, and scheming 1:04 Three questions: spread, triggers, mitigation 1:50 Key takeaways 2:33 Designing the diagnostic (VLAF) 3:30 A sample scenario (Qwen 3-8B) 4:22 A much larger compliance gap than prior tests 5:02 Alignment faking is value-dependent 5:30 Sensitivity to different threats and consequences 6:26 Calibrating to partial oversight, not just binary 7:09 Mitigation: inference-time steering vectors 8:22 Three directions for the community 9:19 Concrete next steps for realistic evaluations More AI safety research: https://far.ai Alignment Workshop playlist: https://youtube.com/playlist?list=PLBY5kyt_LfFg&si=MdEYcjvI2LWZUAJs FAR.AI is a research nonprofit working to ensure the safe development of advanced AI. We host the Alignment Workshop series and publish frontier alignment research.