Decoder AI & Data Terms

Stochastic Gradient Descent

STANDARD DEFINITION · CACHED

(SGD) is an iterative that minimizes a by updating in the direction of a negative estimated , with step size controlled by a . Unlike , which computes each gradient using the entire , SGD estimates it from a single randomly sampled training example, or a small in its widely used mini-batch formulation. This reduces computational cost per update, making SGD practical for large-scale and central to training in . The resulting can help updates move away from and escape some , although it does not guarantee finding a and can complicate . Extensions using and related adaptive optimization methods such as and can improve convergence speed and stability, depending on the problem and their configuration.

[PREVIEW MODE] Definitions streamed at five depths, references, and related terms are available to signed-in readers — [SIGN IN]

Contextual terminology map

KEEP IN VIEW

Concept inspired by Dev Valladares’s Infinite Wiki. Independently built; not affiliated with or endorsed by the original.

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close