Building a Multimodal Model Without Multimodal Training Data

Liquid AI
295 views June 8, 2026

Building a model that handles text, audio, and image usually means training on data that combines all three. Liquid AI's approach is different: train strong text-audio and text-image models separately, then use the shared text space as the bridge. In this discussion, CTO Mathias Lechner talks with Saniya Karwa, who works on multimodal research at Liquid AI. They discuss the core challenge of building a single model that handles text, audio, and image, and how Liquid's architecture addresses discrete and continuous modalities. Saniya explains how the text embedding space is used to align audio and visual representations, meaning the model can perform multimodal tasks even before it's seen mixed-modality training data. Subscribe to follow every conversation: https://www.youtube.com/@liquid-ai-inc Careers at Liquid AI: https://www.liquid.ai/careers

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close