Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Generate 2K video with synchronized stereo audio via the comfyui minimax h3 pipeline — open-weight nodes for text, image, and reference inputs, all in ComfyUI.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Gemini Omni
Gemini Omni Video Generator

Nano Banana2
Best Image Generator
Why Run Video Generation with the comfyui minimax h3 Node Setup
The comfyui minimax h3 node system brings MiniMax's omni-modal model directly into ComfyUI as open weights. It parses text, images, video, and audio in one shared context, then renders synchronized stereo audio alongside the footage — voice, SFX, and music all in a single pass. The output tops out at 2K resolution at 24fps for up to 15 seconds, with every parameter exposed for node-level fine-tuning.
- Audio and Video in SyncVoice, sound effects, and musical score are generated together with the footage and saved in a single MP4 — perfectly aligned from one comfyui minimax h3 pass.
- No Cloud Lock-InBecause the comfyui minimax h3 model runs locally, you control the resolution, clip length, and all diffusion settings directly — with zero dependence on an API.
- Multi-Modal ConditioningMix text prompts, pictures, video sequences, and audio references in a single render to pin down a character, visual style, movement, camera work, or spoken voice using comfyui minimax h3.
Getting Started with the comfyui minimax h3 Workflow
Follow these three steps to generate open-weight video with embedded audio using comfyui minimax h3.
Core Strengths of the comfyui minimax h3 Workflow
A complete local video production stack: three native ComfyUI templates, open-weight omni-modal generation, stereo audio baked in, reference-based control, and optional Sage Attention acceleration — all unified under the comfyui minimax h3 workflow.
Three Ready-Made Workflow Templates
The comfyui minimax h3 template pack includes text-to-video, image-to-video, and reference-to-video presets, each covering a distinct generation mode right out of the box.
One Context for Every Modality
The comfyui minimax h3 model processes text, stills, footage, and audio in a single context, letting you blend all reference types during one generation run.
Reference-Driven Creative Control
Anchor a character's appearance, aesthetic, action, camera movement, or vocal tone using source materials — up to 9 images, 3 video clips, and 3 audio files via the comfyui minimax h3 R2V node.
Clear Text and Brand Rendering
Spelled-out wording and brand marks appear sharply with the comfyui minimax h3 model, while instruction following keeps natural-language descriptions of reference relationships intact.
Sage Attention Performance Boost
Insert the Patch Sage Attention KJ node into the comfyui minimax h3 pipeline to nearly double generation speed while preserving output quality.
Precision Resolution and Duration Grid
The comfyui minimax h3 Resolution Selector derives width and height from aspect ratio and megapixel targets, snapped to the model's 32-multiple canvas and 17-frame-per-block timing at 24fps.
comfyui minimax h3 — Frequently Asked Questions
Quick answers about running the MiniMax H3 model in ComfyUI, covering setup, capabilities, and output settings.
What exactly is the comfyui minimax h3 workflow?
It is ComfyUI's native integration of MiniMax H3, an open-weight general-purpose omni-modal generation model. This workflow produces video with synchronized stereo audio from text, images, video, and audio references in a single forward pass.
What resolution and duration can I get?
The comfyui minimax h3 workflow renders up to 2K resolution at 24fps for roughly 15 seconds. The native canvas is set at a 768px short edge, capped at 768x1344 pixels, and rounded to a multiple of 32.
Which generation modes are available?
The comfyui minimax h3 template library includes three presets: text-to-video (T2V), image-to-video (I2V) with optional first/last frame control, and reference-to-video (R2V) for locking character, style, motion, camera, or voice.
Does it produce actual audio?
Yes — the comfyui minimax h3 model generates native stereo audio covering dialogue, sound effects, and music, all modeled in the same pass and written into a single MP4 file.
How do I begin using it?
Update ComfyUI to version 0.30.0 or later, open Template Library > Video, choose a comfyui minimax h3 preset, and follow the on-screen steps to fetch models from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Is there a way to make generation run faster?
Yes — install SageAttention plus the KJNodes custom nodes, then insert a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow to roughly double the rendering speed.
Launch Your Next Video Project with the comfyui minimax h3 Workflow
Run MiniMax H3 locally in ComfyUI with synchronized stereo audio, open weights, and total parameter freedom — text-to-video, image-to-video, and reference-to-video presets are ready when you are.
