Decoder AI & Data Terms

Quantization

STANDARD DEFINITION · CACHED

is a technique used to reduce the precision of a 's weights and activations to decrease memory footprint and increase inference speed. By mapping high-precision values, typically 32-bit floating-point (), to lower-precision formats such as 8-bit integers (), , or even 4-bit representations, it enables the deployment of on resource-constrained hardware like devices or mobile processors. Common strategies include (PTQ) and (QAT), both aiming to minimize the loss in caused by discretization errors. Major industry players like , , and utilize this technique within frameworks like and to optimize performance on and architectures.

[PREVIEW MODE] Definitions streamed at five depths, references, and related terms are available to signed-in readers — [SIGN IN]

Contextual terminology map

KEEP IN VIEW

Concept inspired by Dev Valladares’s Infinite Wiki. Independently built; not affiliated with or endorsed by the original.

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close