Cerebras Supernova: Matthew Berman on what changes when inference stops being the bottleneck

Cerebras
2,017 views September 9, 2026

What happens when a model finishes so fast that inference is no longer the slowest step in the workflow? At Cerebras Supernova, AI creator Matthew Berman talks through how he tests new models and agents, why he values a short, direct path from prompt to finished work, and what he noticed running GPT-5.6 Sol Ultrafast on Cerebras. When inference sped up, he found he could run far fewer parallel threads and stay just as productive—and the local computer and its tool calls surfaced as the new constraint. 0:00 Why fast, direct AI workflows matter 0:26 Meet Matthew Berman and his path to YouTube 1:14 Building a viral AI audience 1:56 Testing and dogfooding new AI products 3:40 ChatGPT vs. Claude for coding agents 4:33 Model capability vs. agent harness 5:35 Testing GPT-5.6 Sol Ultrafast on Cerebras 6:47 When the CPU becomes the bottleneck Watch the full interview for a builder’s view of how speed reshapes everyday AI work.

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close