Interpretability
Interpretability has multiple competing meanings in and : —understanding model behavior and mechanisms; and —understanding outputs in functional context. In common research usage, simple models and small are often more interpretable than complex systems, including . The distinguishes contextual interpretability from , which concerns representations of the mechanisms underlying a system’s operation. Other literature uses these terms interchangeably or reverses the distinction; explainability is not restricted to , and explanations should faithfully reflect the system’s processes rather than merely justify outputs. These capabilities can support , debugging, risk management, accountability, and compliance, particularly in high-impact settings such as healthcare and finance, but regulatory obligations depend on applicable laws, standards, risk classifications, and context. is broader, also encompassing validity, reliability, safety, security, resilience, transparency, privacy, and fairness with harmful bias managed. and research interpretability to better understand model behavior, investigate risks including , and support ; interpretability alone does not ensure alignment with human intent.
[PREVIEW MODE] Definitions streamed at five depths, references, and related terms are available to signed-in readers — [SIGN IN]
Contextual terminology map
KEEP IN VIEW
Ethics, safety & society
Moving as fast as the field, and easier to overlook
- Explainable AI
- Interpretability
- Model Card
- AI Audit
- Algorithmic Bias
- Disparate Impact
- Digital Divide
- WCAG (Web Content Accessibility Guidelines)
- AI Alignment
- Red Teaming
- EU AI Act
- NIST AI RMF (AI Risk Management Framework)
- Frontier Model
- Deepfake
- Content Credentials
- AI Watermarking
- Job Displacement
- Prompt Injection
- Data Poisoning
- Differential Privacy
- GDPR (General Data Protection Regulation)
- Zero-Day
- End-to-End Encryption
- Data Broker
Infrastructure, markets & the economy
The compute, power and capital behind the boom