PyTorch Day India 2026 Scaling Mixture of Experts for Everyone: High Performance MoE Training in PyT

PyTorch
299 views February 16, 2026

Mixture-of-Experts (MoE) architectures offer a powerful path to scaling model capacity without proportional increases in compute — but training them efficiently at large scale has traditionally required complex distributed systems expertise. This talk walks through how developers can scale MoE models directly in PyTorch, using distributed features and optimized parallelism strategies

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close