Attention Mechanism
An is a component of a that dynamically weights and combines information from different input positions or internal representations, helping capture . In , attention was introduced by Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio in 2014, allowing models to selectively access input representations rather than rely on a single fixed-length vector encoding an entire sequence. Attention can operate over words, image patches, or other elements, but does not by itself solve the in or networks. A prominent variant, , relates positions within the same sequence and is central to the architecture, introduced by researchers at in 2017. Transformer attention uses representations: a compares queries with keys to produce scores, which are normalized into used to compute a of values. By replacing sequential recurrence with attention-based processing, Transformers enable parallel computation across sequence positions during training. They underpin many developed by organizations such as , , and , contributing substantially to advances in .
[PREVIEW MODE] Definitions streamed at five depths, references, and related terms are available to signed-in readers — [SIGN IN]
Contextual terminology map
KEEP IN VIEW
Ethics, safety & society
Moving as fast as the field, and easier to overlook
- Explainable AI
- Interpretability
- Model Card
- AI Audit
- Algorithmic Bias
- Disparate Impact
- Digital Divide
- WCAG (Web Content Accessibility Guidelines)
- AI Alignment
- Red Teaming
- EU AI Act
- NIST AI RMF (AI Risk Management Framework)
- Frontier Model
- Deepfake
- Content Credentials
- AI Watermarking
- Job Displacement
- Prompt Injection
- Data Poisoning
- Differential Privacy
- GDPR (General Data Protection Regulation)
- Zero-Day
- End-to-End Encryption
- Data Broker
Infrastructure, markets & the economy
The compute, power and capital behind the boom