Responsible AI Evaluation ‘’What’’s, and ‘’How’’s of Evaluation

Mila - Institut québécois d'IA
60 views February 10, 2026

Tea Talk January 16, 2026 Standard Machine Learning evaluation techniques overlook systemic harms and important corner cases. Proactive use of techniques borrowed from safety engineering frameworks like FMEA and STPA are essential for anticipating failures and harms. In addition, in the Responsible AI and Fairness communities harms are predominantly reduced to only biases in data, which is only a fraction of harms ML systems can cause and will be addressed in this talk. On the other hand performance of Foundation Models is not uniform across the learned distribution/conditions and could be much worse for underrepresented categories. Therefore, we propose the use of internal model signals, such as local geometry, that can be leveraged to predict model performance in downstream tasks without requiring access to the training data. Finally mitigation techniques to address some of the systemic harms such as concept erasure and unlearning can introduce unintended side effects and should be discussed and addressed

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close