Optimizing Multi-Stage AI Pipelines: Prefix Caching for Faster Autoregressive Inference

PyTorch
359 views September 10, 2026

At PyTorch Conference North America, Ricardo Noriega de Soto, Tech Lead for the vLLM Omni team and Alexander Brooks, Principal Machine Learning Engineer at Red Hat, will demonstrate how extending vLLM’s prefix caching mechanism to multistage pipelines boosts inference speeds while reducing GPU memory overhead. Join us in San Jose on October 20th to learn practical strategies for optimizing complex AI workloads: https://hubs.la/Q04v4SL60

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close