MiniMax H3 Is Here — Everything You NEED to Know

ElevenLabs
7,191 views September 10, 2026

MiniMax H3 and H3 Max just landed in ElevenCreative. Here's everything you need to know. Try MiniMax H3 → https://elevenlabs.io/image-video?utm_source=youtube&utm_medium=organic&utm_campaign=model_release&utm_content=minimax_h3_is_here_everything_you_need_to_know MiniMax H3 is a video model that takes text, images, video and audio as references in a single generation, and generates video with audio, up to 15 seconds at 2K. MiniMax's argument is that video generation got fragmented into separate models for text to video, first and last frame, subject reference and motion reference. H3 is their attempt to handle all of that in one model. In this video we break down how referencing works in practice. You attach your files and tell the model what each one is for in a normal sentence: take the camera movement from this video, use this image for the character, match the voice to this audio clip. The clearer the job you give each reference, the more predictable the result. You can attach up to nine images, three video clips and three audio clips, with a limit of 12 files in total, and each clip has to be between 2 and 15 seconds long. Start and end frames are a separate mode, so you can't combine them with references. We also cover the audio, which is generated alongside the video with dialogue stable in 11 languages, and multi-shot, which is built into the model so you can write timecoded cuts directly into your prompt and get a sequence of scenes with consistent characters, objects and locations from a single generation. Clips run 5 to 15 seconds at 24 frames per second. On resolution, the model generates at 480p or 768p, and the 2K option takes the 768p result and upscales it with your prompt and references still in context, so detail is added rather than guessed. Finally, H3 Max. It's not a bigger model, it's a faster one. MiniMax published H3's model weights, and fal used them to rebuild H3 for speed. The claim is a five second video in about three seconds. The trade-off is resolution, which tops out at 768p. Both models are available now inside ElevenCreative. Head to Image & Video and select MiniMax H3 or MiniMax H3 Max in the model picker. What's covered in the video: • What MiniMax H3 is and the problem it's trying to solve • Referencing video, images and audio in a single prompt • Reference limits, and why start and end frames are a separate mode • Native audio and dialogue in 11 languages • Multi-shot: writing timecoded cuts into your prompt • How the 2K resolution actually works • H3 vs H3 Max: speed, open weights, and the trade-off • How to use both inside ElevenCreative Try MiniMax H3 → https://elevenlabs.io/image-video?utm_source=youtube&utm_medium=organic&utm_campaign=model_release&utm_content=minimax_h3_is_here_everything_you_need_to_know Join the Community • Discord → https://discord.gg/hPE7yT33Qc • Reddit → https://www.reddit.com/r/ElevenLabs/ Links & Resources • ElevenLabs → https://elevenlabs.io/?utm_source=youtube&utm_medium=organic&utm_campaign=model_release&utm_content=minimax_h3_is_here_everything_you_need_to_know • Docs & API → https://elevenlabs.io/docs/?utm_source=youtube&utm_medium=organic&utm_campaign=model_release&utm_content=minimax_h3_is_here_everything_you_need_to_know • Blog → https://elevenlabs.io/blog/?utm_source=youtube&utm_medium=organic&utm_campaign=model_release&utm_content=minimax_h3_is_here_everything_you_need_to_know • Building with the API? Subscribe to ElevenLabs Developers → https://www.youtube.com/channel/UC9yH2AvG2IxvsWQkQoLqL0g Connect with ElevenLabs • Subscribe → https://www.youtube.com/@elevenlabs?sub_confirmation=1 • Instagram → https://www.instagram.com/elevenlabsio • TikTok → https://www.tiktok.com/@elevenlabs • X → https://x.com/elevenlabs Try MiniMax H3 → https://elevenlabs.io/image-video?utm_source=youtube&utm_medium=organic&utm_campaign=model_release&utm_content=minimax_h3_is_here_everything_you_need_to_know About ElevenLabs ElevenLabs is the leading AI voice platform for realistic, context-aware speech generation. Creators, developers, and enterprises use our tools to design voices, dub content, and build expressive AI video and audio workflows.

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close