Prompt Injection
is an application-level security vulnerability and a form of affecting applications and workflows, in which attacker-controlled content attempts to override trusted instructions or redirect system behavior. It can occur directly through crafted user input or indirectly through external content, such as web pages, files, or tool outputs, that the application passes to the model. The underlying problem is inadequate separation of trusted instructions from untrusted data, allowing embedded instructions to be treated as authoritative. Successful attacks may circumvent , override directives in a , cause , generate restricted content, or trigger unauthorized tool use through . Organizations including , , and publish mitigation research and guidance emphasizing layered defenses, such as , attack-detection , restricted tool permissions, and user confirmation for sensitive actions; these measures reduce risk but do not guarantee prevention.
[PREVIEW MODE] Definitions streamed at five depths, references, and related terms are available to signed-in readers — [SIGN IN]
Contextual terminology map
KEEP IN VIEW
Ethics, safety & society
Moving as fast as the field, and easier to overlook
- Explainable AI
- Interpretability
- Model Card
- AI Audit
- Algorithmic Bias
- Disparate Impact
- Digital Divide
- WCAG (Web Content Accessibility Guidelines)
- AI Alignment
- Red Teaming
- EU AI Act
- NIST AI RMF (AI Risk Management Framework)
- Frontier Model
- Deepfake
- Content Credentials
- AI Watermarking
- Job Displacement
- Prompt Injection
- Data Poisoning
- Differential Privacy
- GDPR (General Data Protection Regulation)
- Zero-Day
- End-to-End Encryption
- Data Broker
Infrastructure, markets & the economy
The compute, power and capital behind the boom