Distributed Model Training with Ray at Capital One | Ray Summit 2025
At Ray Summit 2025, Brian Nguyen and Nick Resnick from Capital One share how their team adopted Ray to modernize a complex, multi-framework ML ecosystem—moving from CPU-bound limitations to scalable, GPU-accelerated distributed computing. They begin by outlining Capital One’s distributed compute architecture, where Kubernetes manages cluster resources and the KubeRay Operator deploys and orchestrates GPU-enabled, multi-node Ray clusters. The speakers walk through the technical evolution from a constrained, single-node Ray setup to a robust distributed platform designed to support diverse ML workloads across teams. The session then focuses on a real-world use case: Ray Tune for distributed hyperparameter optimization. After migrating to KubeRay, the team encountered a significant data loading bottleneck that impacted performance. Brian and Nick detail the root causes, the symptoms seen in GPU underutilization and network congestion, and how they approached the debugging process. They conclude with a comparative study of two data loading strategies: Ray Data, Ray’s distributed data ingestion and preprocessing framework A custom manual sharding technique developed in-house They share metrics across memory usage, network I/O, and GPU utilization—showing early results that indicate substantial cost efficiency improvements for large GPU workloads. Attendees will gain practical insights into scaling ML compute with Ray, diagnosing real-world performance bottlenecks, and choosing the right data strategy for high-throughput distributed training. Liked this video? Check out other Ray Summit breakout session recordings https://www.youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh Subscribe to our YouTube channel to stay up-to-date on the future of AI! https://www.youtube.com/c/anyscale 🔗 Connect with us: LinkedIn: https://www.linkedin.com/company/joinanyscale/ X: https://x.com/anyscalecompute Website: https://www.anyscale.com/