Decoder AI & Data Terms

Attention Mechanism

STANDARD DEFINITION · CACHED

An is a component of a that dynamically weights and combines information from different input positions or internal representations, helping capture . In , attention was introduced by Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio in 2014, allowing models to selectively access input representations rather than rely on a single fixed-length vector encoding an entire sequence. Attention can operate over words, image patches, or other elements, but does not by itself solve the in or networks. A prominent variant, , relates positions within the same sequence and is central to the architecture, introduced by researchers at in 2017. Transformer attention uses representations: a compares queries with keys to produce scores, which are normalized into used to compute a of values. By replacing sequential recurrence with attention-based processing, Transformers enable parallel computation across sequence positions during training. They underpin many developed by organizations such as , , and , contributing substantially to advances in .

[PREVIEW MODE] Definitions streamed at five depths, references, and related terms are available to signed-in readers — [SIGN IN]

Contextual terminology map

KEEP IN VIEW

Concept inspired by Dev Valladares’s Infinite Wiki. Independently built; not affiliated with or endorsed by the original.

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close