Cerebras Big Chip Club: Sara Hooker on GPU Bottlenecks in AI
Sara Hooker – researcher, founder of Adaption, and author of the landmark essay "The Hardware Lottery" – sits down with Cerebras to talk about why GPUs are holding AI back, why the hardware lottery has gotten worse since she first wrote about it in 2020, and what the shift to agentic and continuously-learning AI means for the future of compute. Six years ago, Sara argued that hardware, not ideas, determine which AI breakthroughs succeed. In this conversation, she makes the case that we're now in an even more dire moment: the workloads driving the next generation of AI (sequential agentic flows, real-time adaptation, continuous learning, long-tail reasoning) are fundamentally mismatched with the hardware they're running on. Sara's core argument lines up with the reason Cerebras exists: as AI shifts from predictable batch workloads to sequential, memory-bound, adaptive workflows, the costs of moving data between memory and compute become the dominant bottleneck. The Wafer-Scale Engine, with 44GB of on-chip memory and 6000x the memory bandwidth of a GPU, is built for the world Sara describes. 00:00 Introductions & the Hardware Lottery 01:03 Why Adaption? 03:01 Different approaches to continuous learning 06:03 Rethinking AI interfaces to move beyond prompt engineering 08:22 Do AI coding tools worsen the hardware lottery? 12:31 What's changed since the Hardware Lottery essay (2020) 16:26 Motivations behind The Hardware Lottery 18:22 Why GPUs struggle with sparsity 20:28 Why matmul is not the right way to think about data 23:20 General-purpose vs. specialized hardware 27:49 How Adaption co-designs models for GPUs 30:06 How GPU memory limitations affect models 33:23 Cerebras Wafer-Scale Engine and breaking the memory wall 37:04 Can new hardware unlock different model architectures? 38:33 ML techniques held back by today's hardware 41:43 Designing hardware for Adaption 42:42 Predictions for 2026 and beyond +++ Subscribe to our channel! https://www.youtube.com/@Cerebras Cerebras builds the world’s largest AI chip — delivering up to 15× faster inference than leading GPUs. Our mission is to engineer the future of compute and make state-of-the-art AI accessible to every team. Explore our newest open-source models and get free compute at http://cerebras.ai/ . Watch our full video library: https://youtube.com/@Cerebras/videos Read the latest engineering deep dives on our blog: https://cerebras.ai/blog Explore our systems and technology: https://cerebras.ai/publications Follow Cerebras on X: https://x.com/cerebras Connect with us on LinkedIn: https://linkedin.com/company/cerebras-systems