Yonatan Belinkov - Toward Scalable and Actionable Interpretability [Alignment Workshop]

FAR․AI
136 views March 5, 2026

Yonatan Belinkov reveals some actionable findings in interpretability. He shows how we could stop models from makeing basic errors like calling "half a decade" 2.5 years, and how, through targeted interventions, his team can surgically remove sensitive information by projecting neurons into vocabulary space, demoting problematic tokens, and re-embedding them. Belinkov advocates for automation, agents, and core algorithmic improvements to make interpretability actionable for real-world deployment where people use LLMs for complex tasks despite persistent mistakes. Note: The opinions shared in this event are those of the speaker(s) and may not represent the views of FAR.AI or their affiliated organizations.

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close