Load an official template, run it unchanged once, then alter one node at a time.
Creative AI
Build a studio, not a slot machine.
Choose current video, voice and music models, then build a repeatable workflow across hosted tools and ComfyUI.
New coverage · September 8, 2026
Video and voice have moved on.
Start with the current model guides. Each separates documented capabilities, access options and practical evaluation advice.
MiniMax H3.
Seedance 2.5.
Compare hosted video, downloadable H3-Base workflows, references, duration and sound.
Explore video models →Voice, transcription + musicMake it speak.
Give it a soundtrack.
MiniMax Speech 2.8, Cartesia Sonic 3.6, ElevenLabs speech and Scribe, plus MiniMax Music 3.
Explore voice & audio →ComfyUI
The control room for local creative AI.
ComfyUI is a node-based workflow system. You connect model loading, prompting, conditioning, generation, upscaling, and saving into a graph you can reuse or share.
A generated file is one result. The workflow is the durable creative asset.
Community nodes can break or be malicious. Prefer known publishers and lock working versions.
Keep the seed, checkpoint, LoRA, prompt, dimensions, and sampler with approved work.
Local model shelf
More models for your local workflow.
Earlier model shelf, reviewed August 9, 2026. Use the video and voice guides above for newly reviewed releases. Hardware guidance varies with resolution, duration, precision, offloading, and custom nodes.
FLUX.2 [klein] 4B / 9B
Fast text-to-image and editing on creator-class hardware
- Hardware reality
- Start around 12–24GB VRAM depending on precision and workflow
- Know first
- Smaller than FLUX.2 dev and easier to iterate with. Check the exact model license.
FLUX.2 [dev]
High-quality local image generation and more demanding edits
- Hardware reality
- 32B weights; optimized FP8/NVFP4 variants reduce the memory load
- Know first
- The dev weights use a non-commercial / non-production license unless separately licensed.
LTX-2
Text or image to controllable video with synchronized audio
- Hardware reality
- RTX optimization is mature; video remains compute- and storage-heavy
- Know first
- A strong ComfyUI path for storyboard, keyframe, camera, and audio-video experiments.
Stable Audio 3
Music, sound effects, inpainting, and audio continuation
- Hardware reality
- 0.6B small models and 2B medium make local experimentation approachable
- Know first
- Choose Small Music, Small SFX, or Medium; review commercial terms before client work.
A real production workflow
Make a twenty-second product film.
This pipeline keeps creative decisions inspectable and lets you rerun only the stage that needs work.
Define the film
Audience, message, visual references, aspect ratio, duration, required shots, and what must stay true.
Create a look book
Generate still frames first. Approve lighting, materials, character or product consistency, and typography separately.
Animate keyframes
Use image-to-video for each shot. Describe camera, subject action, environment action, and pacing—not the whole film at once.
Build the audio bed
Create music and effects as separate stems. Record or synthesize voice only with permission and clear rights.
Edit like a filmmaker
Cut in a normal editor. Add titles, color, sound mix, captions, provenance, and a human review before publishing.
Hardware reality
Image is approachable. Video eats everything.
Start with shorter, lower-resolution clips. Lock the creative direction with stills before spending minutes or hours on motion.
Learn the workflow
Smaller image models, optimized checkpoints, short video tests, and frequent memory offloading.
Serious creator
Strong image work, FLUX.2 variants, practical LTX-2 pipelines, upscaling, and faster iteration.
Production workstation
Larger models, fewer compromises, higher resolutions, multi-model graphs, and more room for training.
Capacity-first lab
DGX Spark can fit unusually large models; an RTX workstation may still win on speed for common creative graphs.
Earlier cloud pricing snapshot · August 9, 2026
Rent the frontier before you buy the workstation.
Cloud video is ideal for occasional use, comparing frontier quality, or testing whether a concept deserves a local production setup.
Image-to-video; higher per-second price at 720p and 1080p
DetailsOpenAI’s Sora 2, Sora 2 Pro, and Videos API are scheduled to retire September 24, 2026, with no replacement listed. Avoid starting a new integration. Official lifecycle notice
Other creative pricing and hardware guidance was reviewed August 9, 2026. Prices vary by resolution, audio, model tier, and region. API access can differ from consumer product access. Confirm the current provider page before budgeting.