Yonatan Belinkov - Toward Scalable and Actionable Interpretability [Alignment Workshop]
Yonatan Belinkov reveals some actionable findings in interpretability. He shows how we could stop models from makeing basic errors like calling "half a decade" 2.5 years, and how, through targeted interventions, his team can surgically remove sensitive information by projecting neurons into vocabulary space, demoting problematic tokens, and re-embedding them. Belinkov advocates for automation, agents, and core algorithmic improvements to make interpretability actionable for real-world deployment where people use LLMs for complex tasks despite persistent mistakes. Note: The opinions shared in this event are those of the speaker(s) and may not represent the views of FAR.AI or their affiliated organizations.