Optimizing LLM Inference: Disaggregated Serving, PD Protocol, & KV Pinning

PyTorch
661 views September 11, 2026

At PyTorch Conference North America 2026, Nicolò Lucchesi, Research Engineer at Mistral AI and vLLM maintainer, will present joint work with Amazon Web Services (AWS) and Red Hat on how disaggregated serving in vLLM has evolved to support the latest generation of hybrid models. Join us in San Jose on October 20-21: https://hubs.la/Q04v4SL60

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close