What It Takes to Train LFMs at Scale
Training foundation models at scale requires solving parallelism across every layer of the architecture. At Liquid AI, that means working across a hybrid model that combines attention and convolution operators, not a standard transformer stack. In this interview, CTO Mathias Lechner speaks with founding engineer Paul Pak about training infrastructure for Liquid Foundation Models (LFMs): The full parallelism toolkit, and the specific challenge of making context parallelism work correctly across a hybrid model that combines attention and convolution operators. Each requires different handling when the input sequence is sharded across GPUs, and the team built a new algorithm to solve it. If you are interested in distributed systems or ML infrastructure, this episode shows how Liquid AI approaches these problems. Subscribe to follow every episode: https://www.youtube.com/@liquid-ai-inc Careers at Liquid AI: https://www.liquid.ai/careers