Adam Kalai - Consensus Sampling for Safer Generative AI [Alignment Workshop]
Adam Kalai (OpenAI) presents consensus sampling as a cryptography-inspired safety mechanism for generative AI that addresses "provably undetectable harms" from misaligned systems. The approach uses multiple AI models to cross-check outputs by evaluating probability distributions—something "humans can't do." For any unsafe output set, the probability becomes "at most some multiple of the probability of the better of the two models." While effective when one model remains aligned and unsafe outputs are "very, very unlikely," the method requires model overlap and may abstain when models are "totally disjoint." This architecture-agnostic approach "potentially could be used for superintelligent AI." Note: The opinions shared in this event are those of the speaker(s) and may not represent the views of FAR.AI or their affiliated organizations.