Primus Tuning Hybrid Projection and Agentic Search for Distributed Training
How do you train a 671-billion-parameter model across a thousand GPUs, and do it efficiently? It comes down to a lot of choices. How do you split the model across the GPUs? What batch size? How do you schedule the work? Normally, finding the best setup means running option after option, burning weeks and a huge amount of compute. Join Anshu Raina and Peyman Razaghi of AMD at PyTorchCon NA where they show how the Primus Tuning Agent predicts how fast each option will run, ahead of time, from quick hardware benchmarks, so you find the best configuration without running them all, saving thousands of GPU-hours of trial-and-error. Register for PyTorchCon North America today: https://hubs.la/Q04v4SL60