Red Teaming
in is a structured form of in which human testers, automated systems, or both probe an or system to identify vulnerabilities, , misuse risks, and unsafe or prohibited behaviors. Rooted in , it can include attempts to bypass through or , or to elicit and . Used by organizations such as , , and , it supports the of and other AI systems before or after deployment, helping identify safety, security, and issues and inform risk management. Findings can guide mitigations and safety training, including , and may improve model behavior. However, red teaming has limited coverage: it cannot prove the absence of vulnerabilities, guarantee conformity to or robustness standards, or establish with human values.
[PREVIEW MODE] Definitions streamed at five depths, references, and related terms are available to signed-in readers — [SIGN IN]
Contextual terminology map
KEEP IN VIEW
Ethics, safety & society
Moving as fast as the field, and easier to overlook
- Explainable AI
- Interpretability
- Model Card
- AI Audit
- Algorithmic Bias
- Disparate Impact
- Digital Divide
- WCAG (Web Content Accessibility Guidelines)
- AI Alignment
- Red Teaming
- EU AI Act
- NIST AI RMF (AI Risk Management Framework)
- Frontier Model
- Deepfake
- Content Credentials
- AI Watermarking
- Job Displacement
- Prompt Injection
- Data Poisoning
- Differential Privacy
- GDPR (General Data Protection Regulation)
- Zero-Day
- End-to-End Encryption
- Data Broker
Infrastructure, markets & the economy
The compute, power and capital behind the boom