Why Inference—Not Training—Drives AI Infrastructure | WEKA
*What’s driving AI infrastructure in 2025 and beyond?* WEKA Chief AI Officer Val Bercovici sits down with Crusoe’s Kyle Sosnowski to reveal why the industry is shifting from training-focused to inference-optimized workloads, and what memory-bound optimization means for cloud providers and enterprises deploying AI at scale. *Chapters* 00:00 - What Do Subscalers Offer That Hyperscalers Cannot for AI Infrastructure? 01:24 - Why Is This the Decisive Year for Shifting from Training to Inference? 03:33 - How Does Memory Optimization and KV Cache Management Win the Inference Battle? 04:152 - What Is Self-Healing GPU Infrastructure and Why Does It Matter? 07:36 - How Fast Is the AI Application Layer Moving? *Why are companies diversifying away from hyperscalers to subscalers?* The answer lies in specialized GPU infrastructure, intelligent KV cache management, and self-healing systems designed specifically for inference workloads. Kyle discusses how Crusoe's acquisition of Attero positions them to optimize memory-bound workloads, which represent the next frontier in AI infrastructure. *What makes inference different from training?* Inference demands low latency, efficient prompt caching, and automated fault tolerance—capabilities traditional CPU-focused cloud platforms weren't built to handle. The discussion covers how open-source models are democratizing AI application development, why agentic workflows require different optimization strategies than traditional ML pipelines, and how self-healing infrastructure enables enterprises to adopt multi-cloud strategies without increasing operational complexity. 👉 *Hear more from WEKA CTO Shimon Ben-David on how a memory-first architecture accelerates inference:* https://www.weka.io/resources/video/how-memory-first-architecture-solves-ai-inference-challenges/ *What is the subscaler advantage?* Better hardware access, modern features, hands-off management, and infrastructure that "just works" for GPU workloads. *Key topics:* • Memory optimization strategies • The role of KV cache management in reducing inference costs • Automated GPU health monitoring • Why user experience at the infrastructure layer determines competitive advantage in the AI cloud market 👉 *Learn why storage architecture is the new bottleneck impacting scale and capacity:* https://www.weka.io/blog/ai-ml/why-storage-architecture-is-the-new-bottleneck-for-hpc-and-ai-teams/?utm_source=youtube *About the Speaker*: With 20 years of SaaS experience across consumer and enterprise, Kyle led a team in LinkedIn’s Talent Solutions building ML products for SMBs, and now focuses on managed services, enterprise integrations, and the cloud application layer at Crusoe. 👉 *Connect with WEKA:* *Website:* https://www.weka.io?utm_source=youtube&utm_medium=social&utm_campaign=brand *LinkedIn:* https://www.linkedin.com/company/weka-io?utm_source=youtube&utm_medium=social&utm_campaign= *X:* https://x.com/weka?utm_source=youtube&utm_medium=social&utm_campaign=inference