NVIDIA NeMo Curator: Scaling Multi-Modal Data Curation Workflows | Ray Summit 2025
At Ray Summit 2025, Avin Regmi and Matan Appelbaum from Netflix share architectural patterns for processing petabyte-scale, multi-modal data for Generative AI—spanning text, video, audio, and more—using Ray as the backbone for large, heterogeneous distributed pipelines. They begin by outlining why multi-modal data processing at this scale is uniquely challenging. These pipelines must coordinate diverse workloads, manage stateful operations such as deduplication, and efficiently utilize GPU acceleration—all while maintaining reliability across massive distributed clusters. Drawing on their experience building NVIDIA NeMo Curator, the speakers demonstrate how Ray’s core primitives make this possible. They show how: Ray Actors support stateful, long-running operations needed for tasks like deduplication, metadata tracking, and multi-step transformations Ray Tasks provide highly parallel, stateless processing for large-scale batch transformations Heterogeneous CPU/GPU resource management maximizes throughput and pipeline robustness across multi-modal workloads Avin and Matan break down the patterns and best practices that enabled their pipelines to operate efficiently across petabytes of data and thousands of distributed workers. Attendees will walk away with a clear understanding of how to use Ray to architect scalable, resilient, GPU-accelerated data pipelines for next-generation Generative AI applications. Liked this video? Check out other Ray Summit breakout session recordings https://www.youtube.com/playlist?list=PLzTswPQNepXllnU0C36WtkC0dqkAoDulh Subscribe to our YouTube channel to stay up-to-date on the future of AI! https://www.youtube.com/c/anyscale 🔗 Connect with us: LinkedIn: https://www.linkedin.com/company/joinanyscale/ X: https://x.com/anyscalecompute Website: https://www.anyscale.com/