Tomek Korbak - Chain of Thought Monitorability for AI Safety [Alignment Workshop]
Tomek Korbak (OpenAI) reveals transformer architecture creates an AI safety opportunity through chain-of-thought monitoring. We got lucky: models reason in natural language rather than opaque vector spaces, making their cognition readable. Two factors enable monitoring: transformers must externalize reasoning for difficult tasks due to architectural limitations, and models struggle to suppress information even when instructed. However, four threats could close this window: drift from legible English, optimization pressure, indirect optimization effects, and shifts toward latent space reasoning. Korbak recommends developing monitorability measures for training and deployment decisions while this transparency window remains open. Note: The opinions shared in this event are those of the speaker(s) and may not represent the views of FAR.AI or their affiliated organizations.