Building a Multimodal Model Without Multimodal Training Data
Building a model that handles text, audio, and image usually means training on data that combines all three. Liquid AI's approach is different: train strong text-audio and text-image models separately, then use the shared text space as the bridge. In this discussion, CTO Mathias Lechner talks with Saniya Karwa, who works on multimodal research at Liquid AI. They discuss the core challenge of building a single model that handles text, audio, and image, and how Liquid's architecture addresses discrete and continuous modalities. Saniya explains how the text embedding space is used to align audio and visual representations, meaning the model can perform multimodal tasks even before it's seen mixed-modality training data. Subscribe to follow every conversation: https://www.youtube.com/@liquid-ai-inc Careers at Liquid AI: https://www.liquid.ai/careers