Cerebras Supernova: Mostafa Elhoushi (Cerebras Core ML) on smaller, smarter models

Cerebras
854 views September 10, 2026

What if every token did not need every layer? At Cerebras Supernova, Cerebras Research Scientist Mostafa Elhoushi sits down with Alycia Cary to revisit dropout for modern language models. Mostafa explains why the technique fell out of favor as training datasets grew, what his team found by applying different dropout rates across layers, and how layer dropout can make models more robust to exiting early at inference. The conversation moves from dropout to sparsity and mixture of experts, then to a larger question: what comes after scaling? Mostafa shares a vision for smaller, smarter architectures that separate knowledge from reasoning—and explains why fast iteration on wafer-scale systems gives researchers more chances to find the idea that works. 0:00 Why dropout disappeared from LLM training 1:34 Why different layers need different dropout rates 2:35 Hyperparameters that make dropout work 3:26 Scaling layer dropout to larger models 4:51 99% dropout and 25% fewer training FLOPs 5:36 Layer skipping and early exit at inference 6:44 A human analogy for interruptible models 7:49 Why not every token needs every layer 9:10 Layer dropout vs. mixture of experts 10:22 Hardware–software co-design at Cerebras 11:28 What comes after model scaling? 13:14 Wafer-scale speed and research iteration Read the paper: https://arxiv.org/abs/2609.05275 Watch the full interview for a technical look at the co-design research happening at Cerebras.

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close