Enterprise gen AI inference demo with vLLM on Red Hat AI stack

Red Hat
193 views August 11, 2026

How can you scale enterprise LLM inference to production? In this demo from Red Hat Summit, Grace Ableidinger, AI Developer Advocate at Red Hat, breaks down how to scale AI workloads from a single RAG chatbot to a multi-model enterprise architecture. Learn how AI Gateway manages control plane routing, how vLLM handles paged attention and continuous batching, and how llm-d optimizes routing across replicas. Discover optimization techniques including model quantization, speculative decoding, and prefill-decode disaggregation to reduce latency and maximize GPU efficiency. 00:00 Introduction to enterprise gen AI inference 00:15 AI Gateway, vLLM, and llm-d architecture overview 00:49 Scaling a single GPU RAG chatbot with vLLM 01:16 Model quantization with vLLM and Hugging Face 01:50 Speculative decoding for faster token generation 02:20 Horizontal scaling and smart routing with llm-d 03:34 Prefill-decode disaggregation for agentic workflows 04:06 Exploring the Red Hat AI model catalog Find more gen AI resources from Red Hat: ✨ Explore Red Hat AI → https://www.redhat.com/en/products/ai 📖 Read the Red Hat AI blog → https://www.redhat.com/en/blog/channel/artificial-intelligence 🤝 Browse models on Red Hat AI Hugging Face → https://huggingface.co/RedHat #RedHatAI #OpenShiftAI #vLLM #LLMD #GenAI #AIInference #LLM #MachineLearning

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close