Enterprise gen AI inference demo with vLLM on Red Hat AI stack
How can you scale enterprise LLM inference to production? In this demo from Red Hat Summit, Grace Ableidinger, AI Developer Advocate at Red Hat, breaks down how to scale AI workloads from a single RAG chatbot to a multi-model enterprise architecture. Learn how AI Gateway manages control plane routing, how vLLM handles paged attention and continuous batching, and how llm-d optimizes routing across replicas. Discover optimization techniques including model quantization, speculative decoding, and prefill-decode disaggregation to reduce latency and maximize GPU efficiency. 00:00 Introduction to enterprise gen AI inference 00:15 AI Gateway, vLLM, and llm-d architecture overview 00:49 Scaling a single GPU RAG chatbot with vLLM 01:16 Model quantization with vLLM and Hugging Face 01:50 Speculative decoding for faster token generation 02:20 Horizontal scaling and smart routing with llm-d 03:34 Prefill-decode disaggregation for agentic workflows 04:06 Exploring the Red Hat AI model catalog Find more gen AI resources from Red Hat: ✨ Explore Red Hat AI → https://www.redhat.com/en/products/ai 📖 Read the Red Hat AI blog → https://www.redhat.com/en/blog/channel/artificial-intelligence 🤝 Browse models on Red Hat AI Hugging Face → https://huggingface.co/RedHat #RedHatAI #OpenShiftAI #vLLM #LLMD #GenAI #AIInference #LLM #MachineLearning