Scaling Multi-Modal Datasets to Petabytes with Ray at Apple | Ray Summit 2025
At Ray Summit 2025, Yufan Fei, Haocheng Bian, and Andy Wu from Apple share how they scaled multimodal data processing frameworks to handle petabyte-scale workloads as LLM applications grow increasingly complex. They begin by outlining the rising demand for large-scale processing over multimodal datasets and the challenges of designing systems that are both user-friendly and capable of operating at extreme scale. While Ray’s distributed computing model offers strong flexibility, achieving petabyte-level scalability requires addressing bottlenecks across the data processing layer, Ray’s core runtime, and the underlying infrastructure. The speakers then dive into the key lessons learned from scaling multimodal workloads to this magnitude. They discuss strategies for mitigating performance bottlenecks, ensuring infrastructure resilience, and tuning Ray’s internal components to prevent cascading failures that can arise in massive, tightly coupled distributed systems. Finally, they share practical takeaways from architecting and productionizing Apple’s petabyte-scale multimodal pipeline as a self-serve service, balancing usability with the ability to process enormous datasets reliably and efficiently. Attendees will walk away with actionable insights for building and operating large-scale multimodal data workflows in real-world production environments. Subscribe to our YouTube channel to stay up-to-date on the future of AI! https://www.youtube.com/c/anyscale 🔗 Connect with us: LinkedIn: https://www.linkedin.com/company/joinanyscale/ X: https://x.com/anyscalecompute Website: https://www.anyscale.com/