Inference
Running models in production: serving stacks, latency, throughput and the engines behind them
Latest stories See all →
- Still Batching Streaming Data into Files? Send Kafka Straight to the Lake with Kafka Connect [OSS Tables Deep Dive]
- OpenAI’s voice model doesn’t think. That’s the point.
- Inference Cost Optimization: How to Cut Your AI Bill by 75%
- Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train
- How NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera Rubin
- Your Inference Server is Secretly a Learner: Reef Infrastructure for Continual Self-Improving Agents
- The AI Inference Revolution Is Here
- AxeleraScript: One Language, every Inference workload solved
- Introducing d-Matrix Demo Cloud: Ultra Low Latency Inference, Ready to Test in Minutes
- Scaling ML Model Inference through architectural separation
- Best Open Source LLM API Providers in 2026
- Which GPU should you actually use for embedding workloads?
- 9/14/2026DeepSeek-V4.1-Flash on Fireworks: Astra-level DeepSWE at 1/15th the cost
- Breaking Through the KV Cache Memory Wall with DPU-powered Network Storage
- The Cost of Keeping a Model Catalog Current
- vLLM x Novita AI: Chord, Faster INT4 MoE for Kimi K2.x. Up to 1.3x on H200, 2.15x on Untuned B300
- Vera Rubin NVL72 Agentic Inference: 67x better Performance per Dollar
- A Brain Too Big to Carry — On-Device vs Datacenter Inference
- Run 40% more post-training experiments on the same GPUs with llm-d time-slicing
- DeepSeek-V4.1-Flash Packs 552B Parameters With Efficient MoE Inference
- Speculative Decoding in Llama.cpp: How to Actually Speed Up Local LLMs
- How to Run OUI-1 with vLLM for Generative UI
- Chip Huyen explains how to cut inference costs without new hardware
- Deterministic LLM Inference for Gemma 4 on Windows XP
Podcasts
Recent episodes
- Apple’s Ternus Era Starts the iPhone Duo
- Why the intelligence explosion can't happen inside a data centre | Tom Reed
- Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776
- Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776
- The Five-Layer Cake Approach to Scaling AI Without Wasting Money
- The Death of Online Anonymity
- The Death of Online Anonymity
- The Death of Online Anonymity
- The Death of Online Anonymity
- EP 190: NVIDIA and Marvell Earnings, Hot Chips Hot Takes
- #255 - Gemini 3.7, Jalapeño, Qwen 3.8, Drones
- OpenAI’s Jalapeño chip is built for fast inference at scale; plus, Apple debuts its ‘most powerful chip ever’ in M5 Ultra and M6