Building Fault-Tolerant Massive Ray Clusters on Anyscale | Ray Summit 2025
Slides: https://drive.google.com/file/d/1rnSSECmBeigaZDOxnbsq8GSh2mfZep55/view?usp=sharing At Ray Summit 2025, Dhyey Shah and Ibrahim Rabbani from Anyscale share how Ray is engineered to withstand the harsh realities of large-scale AI workloads—and what it takes to run reliably on clusters exceeding 10,000 nodes. They begin by outlining the challenges that inevitably arise at extreme scale: network flakiness, spot preemptions, hardware failures, resource contention, and unpredictable infrastructure behavior. The speakers explain how Ray’s distributed runtime is designed to absorb these failures gracefully, keep workloads running, and maintain strong reliability even under constant churn. Dhyey and Ibrahim then dive into key engineering lessons learned from building and operating Ray at massive scale, including strategies for fault tolerance, state management, recovery, elasticity, and workload-aware scheduling. Finally, they share what’s next for Ray—highlighting upcoming improvements to scalability, reliability, and performance aimed at powering the next generation of AI applications. Whether you're running large-scale model training, distributed inference, reinforcement learning, or multimodal pipelines, this session offers a deep look at how Ray stays resilient when everything else breaks. Liked this video? Check out other Ray Summit breakout session recordings https://www.youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh Subscribe to our YouTube channel to stay up-to-date on the future of AI! https://www.youtube.com/c/anyscale 🔗 Connect with us: LinkedIn: https://www.linkedin.com/company/joinanyscale/ X: https://x.com/anyscalecompute Website: https://www.anyscale.com/