AI Alignment
is a subfield of concerned with making the goals, behaviors, and decision-making processes of systems consistent with human intentions, values, and ethical principles. A common distinction separates , which asks whether a specified or matches designers’ intended goals, from , which asks whether a model’s learned objective matches its specified training objective. Organizations including , , and research alignment and safety for and prospective ; examples of their approaches include , used by , and , developed by . These training and mitigation methods have known limitations and do not guarantee prevention of or ensure . Alignment research can overlap with mitigation and , but it is broader than and distinct from those areas, and does not imply eliminating bias. It also investigates potential associated with , including theoretical concerns about , rather than treating those risks as established properties of deployed systems.
[PREVIEW MODE] Definitions streamed at five depths, references, and related terms are available to signed-in readers — [SIGN IN]
Contextual terminology map
KEEP IN VIEW
Ethics, safety & society
Moving as fast as the field, and easier to overlook
- Explainable AI
- Interpretability
- Model Card
- AI Audit
- Algorithmic Bias
- Disparate Impact
- Digital Divide
- WCAG (Web Content Accessibility Guidelines)
- AI Alignment
- Red Teaming
- EU AI Act
- NIST AI RMF (AI Risk Management Framework)
- Frontier Model
- Deepfake
- Content Credentials
- AI Watermarking
- Job Displacement
- Prompt Injection
- Data Poisoning
- Differential Privacy
- GDPR (General Data Protection Regulation)
- Zero-Day
- End-to-End Encryption
- Data Broker
Infrastructure, markets & the economy
The compute, power and capital behind the boom