Chenhao Tan - Automating Mechanistic Interpretability [Alignment Workshop]

FAR․AI
548 views March 5, 2026

Chenhao Tan demonstrates an automated mechanistic interpretability agent, MechEvalAgent, and describes the current state of this work-in-progress. He argues that mechanistic interpretability represents "a dream problem for research agents" because experiments run in silico with causally testable findings. AI models like Claude and Gemini already conduct interpretability research when instructed, but in 2025, the bottleneck shifted from agents running experiments, to reviewing and trusting these agents. His MechEvalAgent framework validates automated mechanistic interpretability research using three dimensions: coherence (following plans), reproducibility (rerunning experiments), and generalizability (extending findings). The research exposes critical failure modes for automated evaluation agents, including hallucinations where agents claim ablation studies but code shows otherwise. He presents a core challenge that automation will not advance interpretability unless we automate evaluation to some extent, and yet, we do not yet have the tools to evaluate automated evaluation agents. Note: The opinions shared in this event are those of the speaker(s) and may not represent the views of FAR.AI or their affiliated organizations.

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close