Christopher Summerfield - Lessons from a Chimp: AI "Scheming" & the Quest for Ape Language [Alignmen
Christopher Summerfield, Professor of Cognitive Neuroscience at Oxford and Research Director at the UK AI Safety Institute (AISI), argues that research into AI model intentionality needs much stronger experimental rigor. Drawing a parallel with the discredited 1970s research into ape language — where researchers inadvertently prompted the very behaviors they were looking for — he warns that some current AI scheming research follows a similar pattern: iterating scenarios until misaligned behavior appears, then reporting those incidents without adequate controls. This raises the possibility that methodological circularity in current research may be distorting our picture of how often misaligned behavior actually occurs. As a preliminary example of more rigorous methodology, he presents work from AISI examining whether LLMs' stable charity preferences predict sandbagging behavior in a prize quiz task — finding that while models show stable preferences, there is limited evidence they exploit those preferences to act in misaligned ways. Note: The opinions shared in this event are those of the speaker(s) and may not represent the views of FAR.AI or their affiliated organizations.