Topics

Inference

Running models in production: serving stacks, latency, throughput and the engines behind them

271 stories in 30 days · 107 sources · Filter the feed

Latest stories See all →

  1. Still Batching Streaming Data into Files? Send Kafka Straight to the Lake with Kafka Connect [OSS Tables Deep Dive]
  2. OpenAI’s voice model doesn’t think. That’s the point.
  3. Inference Cost Optimization: How to Cut Your AI Bill by 75%
  4. Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train
  5. How NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera Rubin
  6. Your Inference Server is Secretly a Learner: Reef Infrastructure for Continual Self-Improving Agents
  7. The AI Inference Revolution Is Here
  8. AxeleraScript: One Language, every Inference workload solved
  9. Introducing d-Matrix Demo Cloud: Ultra Low Latency Inference, Ready to Test in Minutes
  10. Scaling ML Model Inference through architectural separation
  11. Best Open Source LLM API Providers in 2026
  12. Which GPU should you actually use for embedding workloads?
  13. 9/14/2026DeepSeek-V4.1-Flash on Fireworks: Astra-level DeepSWE at 1/15th the cost
  14. Breaking Through the KV Cache Memory Wall with DPU-powered Network Storage
  15. The Cost of Keeping a Model Catalog Current
  16. vLLM x Novita AI: Chord, Faster INT4 MoE for Kimi K2.x. Up to 1.3x on H200, 2.15x on Untuned B300
  17. Vera Rubin NVL72 Agentic Inference: 67x better Performance per Dollar
  18. A Brain Too Big to Carry — On-Device vs Datacenter Inference
  19. Run 40% more post-training experiments on the same GPUs with llm-d time-slicing
  20. DeepSeek-V4.1-Flash Packs 552B Parameters With Efficient MoE Inference
  21. Speculative Decoding in Llama.cpp: How to Actually Speed Up Local LLMs
  22. How to Run OUI-1 with vLLM for Generative UI
  23. Chip Huyen explains how to cut inference costs without new hardware
  24. Deterministic LLM Inference for Gemma 4 on Windows XP

Podcasts

Recent episodes

  1. Apple’s Ternus Era Starts the iPhone Duo
  2. Why the intelligence explosion can't happen inside a data centre | Tom Reed
  3. Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776
  4. Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776
  5. The Five-Layer Cake Approach to Scaling AI Without Wasting Money
  6. The Death of Online Anonymity
  7. The Death of Online Anonymity
  8. The Death of Online Anonymity
  9. The Death of Online Anonymity
  10. EP 190: NVIDIA and Marvell Earnings, Hot Chips Hot Takes
  11. #255 - Gemini 3.7, Jalapeño, Qwen 3.8, Drones
  12. OpenAI’s Jalapeño chip is built for fast inference at scale; plus, Apple debuts its ‘most powerful chip ever’ in M5 Ultra and M6

From sources that build this

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close