Responsible AI Evaluation ‘’What’’s, and ‘’How’’s of Evaluation
Tea Talk January 16, 2026 Standard Machine Learning evaluation techniques overlook systemic harms and important corner cases. Proactive use of techniques borrowed from safety engineering frameworks like FMEA and STPA are essential for anticipating failures and harms. In addition, in the Responsible AI and Fairness communities harms are predominantly reduced to only biases in data, which is only a fraction of harms ML systems can cause and will be addressed in this talk. On the other hand performance of Foundation Models is not uniform across the learned distribution/conditions and could be much worse for underrepresented categories. Therefore, we propose the use of internal model signals, such as local geometry, that can be leveraged to predict model performance in downstream tasks without requiring access to the training data. Finally mitigation techniques to address some of the systemic harms such as concept erasure and unlearning can introduce unintended side effects and should be discussed and addressed