Causal Representation Learning: A Natural Fit for Mechanistic Interpretability

Mila - Institut québécois d'IA
84 views February 10, 2026

Tea Talk November 28, 2025 As the capabilities of large language models (LLMs) grow, so too does the need to interpret the internal mechanisms of these models, which can enable us to steer model behavior. Interpretability typically requires supervision from e.g., contrastive pairs of prompts that vary by a single target concept, which is costly to obtain and limits the speed of research. An appealing alternative is to use unsupervised approaches such as sparse autoencoders (SAEs) to map LLM embeddings to sparse representations that capture human-interpretable concepts. However, without further assumptions, SAEs may not be identifiable: they could learn latent dimensions that entangle multiple concepts, leading to incorrect understanding and unintentional steering of properties. In this talk, I'll introduce sparse shift autoencoders (SSAEs), identifiable models inspired by causal representation learning. These models map the differences between embeddings to sparse representations that capture concept shifts. Crucially, we show that SSAEs are identifiable from paired observations that vary by multiple unknown concepts, leading to accurate steering of single concepts without the need for supervision. We empirically demonstrate accurate steering across semi-synthetic and real-world language datasets.

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close