AI Token Economics and Prompt Caching Optimization | SemiAnalysis x WEKA
*How do AI token economics work?* Why does prompt caching matter for reducing inference costs? Wei Zhou from SemiAnalysis and Val Bercovici from WEKA discuss the evolution of token pricing, KV cache hit rates, and GPU benchmarking strategies at AI Infra Summit 2025. *What is token economics in AI?* Token economics refers to the cost structure of serving AI models, where software has shifted from zero marginal costs to relatively high marginal costs. The price of input tokens, output tokens, and cached tokens reflects the hardware utilization required to produce each token. As Wei Zhou explains, "You can almost print any number you want—you want to produce a two-cent token, you can produce a two-cent token; you want to produce a $10 token, you produce a $10 token." The economic trade off depends on tokens per user per second and time to first token. Why does prompt caching reduce AI costs? Wei and Val discuss how prompt caching can be a price discriminator allowing users to get cached tokens at a 90% discount or essentially free with owned hardware. For agentic workloads with 10:1, 50:1, or even 100:1 input-to-output ratios, KV cache hit rates become the number one indicator of agent success. As context windows expand to millions of tokens, prefix cache hits eliminate the need to re-prefill expensive input tokens. *Chapters* 0:00 - What Are Token Economics in AI? How Software Shifted from Zero to High Marginal Costs 1:03 - Who Pays for AI Tokens? Breaking Down GPU Providers, Model Builders, and Enterprise Users 3:06 - How Does Prompt Caching Reduce AI Costs? Comparing Anthropic and OpenAI's Pricing Models 5:05 - What If Prompt Caching Lasted Days Instead of Hours? Extended Cache Benefits for Agent Swarms 7:13 - Why Is KV Cache Hit Rate the #1 Metric for AI Agent Success? 50% vs 90% Performance Impact 8:47 - How Do Reinforcement Learning Costs Scale in 2025? Memory-Bound Inference vs Pre-Training 11:03 - What Is SemiAnalysis Inference Max? Benchmarking B200, H200, and MI350 GPU Performance 13:15 - How to Model Real AI Workloads: 8:1 vs 100:1 Input-Output Ratios in Production Systems 15:27 - Can AI Benchmarking Follow the Kubernetes Foundation Model? $50K/Week Infrastructure Costs *What is SemiAnalysis Inference Max?* Inference Max benchmarks GPU performance across models like GPT-4-OSS-120B and DeepSeek-R1 on accelerators including NVIDIA B200, H200, and AMD MI350. The platform measures unit economics—tokens per second per user vs. GPU throughput—revealing the economic tradeoffs between user experience and hardware efficiency. 👉 *Explore WEKA's storage solutions for AI inference:* https://www.weka.io/lp/break-the-inference-barrier/?utm_source=youtube&utm_medium=social&utm_campaign=inference *Key Topics Covered:* • Token pricing models: input vs. output vs. cached tokens • Prompt caching strategies for Anthropic and OpenAI APIs • KV cache hit rate optimization (50% to 90% improvement) • Reinforcement learning as memory-bound inference workloads • Real-world input-output ratios (8:1 vs. 100:1 in production) • GPU benchmarking methodology and infrastructure costs ($50K/week) • Building industry standards like the Kubernetes Foundation model *About the Speakers:* Wei Zhou is head of AI utility research at SemiAnalysis focusing on AI infrastructure from hardware to tokens. Val Bercovici is an AI Strategist at WEKA, specializing in high-performance storage for AI/ML workloads. 👉 *Connect with WEKA:* *Website:* https://www.weka.io?utm_source=youtube&utm_medium=social&utm_campaign=brand *LinkedIn:* https://www.linkedin.com/company/weka-io?utm_source=youtube&utm_medium=social&utm_campaign= *X:* https://x.com/weka?utm_source=youtube&utm_medium=social&utm_campaign=inference #AIInfrastructure #TokenEconomics #AIInference #MachineLearning #WEKA