Video models · September 2026
Choose the model behind the motion.
Start with MiniMax H3 and Seedance 2.5, then compare the same scene, references, and sound requirements. Pick an access path before committing to a workflow.
Current shortlist
Three paths into video generation.
These are documented capabilities and editorial starting points, not results from our own head-to-head benchmark. Each card links to the source behind its access and limits.
Hosted · VideoMiniMax H3 · hosted
Generate video with synchronized sound from text and references.
Where to use it
MiniMax API / Hailuo app
4–15 seconds. Full hosted system supports output up to 2K; that is not the downloadable base checkpoint’s resolution.
Model identifier
MiniMax-H3 Downloadable · VideoMiniMax H3-Base
Build a local video-and-audio workflow with downloadable weights.
Where to use it
Official weights / ComfyUI
768p base output. FL2VA uses text and first/last frames; Ref2VA accepts multimodal references. Context-IR and Regenerate-2K are not released. Check the community license and runtime requirements.
Model identifier
MiniMaxAI/MiniMax-H3 · FL2VA / Ref2VA Hosted · VideoSeedance 2.5
Reference-led scenes, longer takes, and targeted video editing.
Where to use it
BytePlus LAS API / Jimeng / Doubao Pro
Up to 30 seconds with multi-round extension. The API ID shown belongs to BytePlus LAS. Consumer rollout, account access and API billing are separate.
Model identifier
dreamina-seedance-2-5-260628 Make the choice
What matters for your film?
Seedance is listed here as a hosted service. A remote model used inside a creative tool still sends the job to a provider; that does not establish downloadable weights or offline operation.
Run a useful comparison in four shots
Use a fictional product or footage you have permission to use. Prepare one reference image, one aspect ratio and a short brief. Generate a product close-up, a simple camera move, a person interacting with the product, and a short spoken scene. Give each candidate the same creative requirements while respecting its input format.
- Save the setup. Record model and version, references, prompt, duration, resolution, sound settings and seed where available.
- Review usable footage. Score identity consistency, motion, physical coherence, sound timing and editability. Count rejected generations.
- Calculate finished cost. Divide total generation and retry spend by approved seconds. Include extensions, upscaling and finishing.
- Choose per shot. One model may suit dialogue while another fits movement or references. Keep the approved stills and final edit as the common foundation.
A first local H3 experiment
Open the official ComfyUI tutorial, choose the text/image or reference workflow, and follow its matching checkpoint and encoder requirements. Run the supplied workflow once before changing it. Start with a short base-resolution result; increase complexity only after measuring your setup.
Official H3 ComfyUI workflows ↗
Before you budget a production
Confirm the provider’s current resolution, duration, sound, concurrency and regional limits in your account. Compare credits or per-second charges against the output you need. Text-token prices from the language-model calculator do not price video.
Choose narration, transcription and music →
Return to the full creative workflow →