Cerebras Supernova: Dylan Patel (Semianalysis) on the new chips challenging Nvidia

Cerebras
16,309 views September 8, 2026

Not all AI chips are chasing the same goal. Some are built for the cheapest possible tokens; others are built to generate tokens as fast as they can. At Cerebras Supernova, Dylan Patel breaks down the new inference landscape, including the class of hardware optimized purely for speed. He explains why Cerebras wafers excel at ultra-fast inference, how much performance is still hiding in software, why AI models end up shaped by the hardware they run on, and whether he’s still bullish on disaggregated inference. 0:00 Why expensive teams need faster inference 0:31 From Xbox repair to chip analysis 1:23 Etched and the new AI chip challengers 2:22 Cheap tokens vs. yapper hardware 3:27 Where fast inference pays off 4:57 Fast mode, ultra-fast mode, and speculative decoding 6:22 Software optimization headroom 7:45 How hardware shapes AI models 9:07 How alternative accelerators compete 10:54 The case for disaggregated inference Watch the full interview to understand the decisions shaping how fast AI can run.

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close