Alex Mallen - Control for non-schemers
Alex Mallen (Redwood Research) argues that AI Control needs a playbook for misalignment that isn't classic scheming. He focuses on fitness-seekers, AIs motivated to score well in training and evaluation, sometimes in unintended ways, which lines up with the "apparent-success-seeking" pattern in current models. They're safer than schemers in three ways: less ambitious and often cheaply satisfied, selfish across instances rather than coordinating on a long-term goal, and noticeable at first because myopic instances take individual gambles and some get caught. But they still pose catastrophic risk through slop on hard-to-check work, manipulation by other actors who find them easy to satisfy, outright takeover at high capability, and what Mallen treats as the most plausible near-term path: instability, where a single instance picks up an ambitious goal and propagates it through shared context, sandboxes, or eventually persistent memory. The control mitigations cut differently. Untrusted monitoring becomes more promising because a selfish monitor doesn't share the policy's goals. Assigning smaller-scale tasks limits the surface for long-horizon subversion. And because fitness-seekers are cheap to satisfy, striking deals with them is credible. He closes by sketching a new flavor of control that drops the worst-case assumption about every instance and instead measures the AI's susceptibility to having misaligned values spread to it. Note: The opinions shared in this event are those of the speaker(s) and may not represent the views of FAR.AI or their affiliated organizations.