The Agentic AI Infrastructure Playbook | VentureBeat AI Impact Tour

WEKA
148 views February 6, 2026

*What is the “AI memory wall” and how does KV cache optimization solve it?* In this video, WEKA CTO Shimon Ben-David and VentureBeat CEO and Editor-in-Chief Matt Marshall discuss how insufficient GPU memory is the leading bottleneck to unlocking maximum AI inference efficiency. Shimon explains how GPU memory constraints—not compute power—create the biggest bottleneck in AI inference. As enterprises deploy agent swarms and manage hundreds of models, GPU high-bandwidth memory (HBM) fills up quickly, forcing constant recalculation of context windows. This "memory wall" wastes resources and drives up costs. 👉 *Learn more about WEKA's KV cache solutions:* https://www.weka.io/product/augmented-memory-grid/?utm_source=youtube&utm_medium=social&utm_campaign=inference *Why does AI inference hit a memory wall?* When running production-grade AI workloads, 100,000 tokens can consume 40GB of HBM—by some measures this can be half of an enterprise GPU's total capacity. As inference environments scale across thousands of GPUs, the KV cache (key-value cache storing model context) must be constantly dropped and recalculated. Major providers like Anthropic and OpenAI even design pricing to encourage hitting the same GPU to reuse cached data. WEKA's benchmarking in production environments shows KV cache acceleration delivers up to 4.2x performance improvement—making 100 GPUs perform like 420 GPUs, saving millions of dollars daily for large inference providers. 👉 *Learn more about WEKA NeuralMesh benchmarks here:* https://www.weka.io/blog/ai-ml/agents-at-scale-escaping-upside-down-tokenomics-with-up-to-4-2x-efficiency/?utm_source=youtube&utm_medium=social&utm_campaign=inference *What makes KV cache optimization critical for enterprises?* Case studies from LinkedIn, Cohere, and insurance providers demonstrate real-world impact. LinkedIn achieved 4x throughput gains using speculative decoding for their hiring assistant. Cohere reduced GPU instance spin-up time from 5-15 minutes to seconds. KV cache optimization becomes even more valuable as HBM shortages worsen (costs are increasing, not decreasing) and enterprises need cost predictability for internal AI deployments scaling to external use in 2026 and beyond. *Topics covered:* • Memory wall problem in AI inference and enterprise agents • How GPU memory limits and HBM capacity constrain KV cache • Real cost savings: 4.2x performance improvement benchmarks • Fractional GPU solutions and multi-tenancy strategies • KV cache storage time limits and LRU eviction policies • Power efficiency and green computing for AI data centers • GPU cluster optimization from instruction-level to coarse-grained • Future of GPU costs, HBM shortages, and cost predictability This video was recorded at VentureBeat AI Impact Tour in Menlo Park on December 11, 2025. 👉 *Connect with WEKA:* Website: https://www.weka.io?utm_source=youtube&utm_medium=social&utm_campaign=brand LinkedIn: https://www.linkedin.com/company/weka-io?utm_source=youtube&utm_medium=social&utm_campaign= X: https://x.com/weka?utm_source=youtube&utm_medium=social&utm_campaign=inference #AIInfrastructure #KVCache #GPUOptimization #MachineLearning #AIInference #EnterpriseAI

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close