MiniMax H3 REF2VA + ComfyUI
Deploy MiniMax H3 REF2VA + ComfyUI on a dedicated GPU from $0.45 per hour, billed per minute, on 1× RTX 4090. MiniMax H3's reference-conditioned audio-video checkpoint, built to generate new scenes from text plus image, video and audio references while preserving subject, style, motion and voice cues.
HexGrid Cloud price
$0.45/hr
1× RTX 4090
Parameters
33B
dense Omni Transformer
Output
768p
local H3-Base stage
Duration
4–15s
24 FPS
Running MiniMax H3 REF2VA + ComfyUI on HexGrid Cloud
What you get when you deploy with us, beyond the hourly rate.
Single tenant by default
The model runs on a GPU that is yours for the life of the instance. No shared endpoint, no queue behind other tenants.
Per-minute billing
You pay for the GPU, not the token. Stop the instance and billing stops with it.
OpenAI-compatible endpoint
Point an existing SDK at your instance by changing the base URL. No rewrite required.
Your weights, your data
Prompts and outputs stay inside your instance. Nothing is logged, sampled or used for training.
MiniMax H3 REF2VA + ComfyUI GPU sizing and cost
What the model needs at each precision, and the cheapest HexGrid Cloud configuration that holds it.
| Precision | vRAM needed | Runs on | Price |
|---|---|---|---|
BF16 Recommended | 80 GB | 5× RTX A4000 | $1.00/hr |
FP8 | 40 GB | 1× RTX A6000 | $0.55/hr |
INT4 | 20 GB | 1× RTX 4090 | $0.45/hr |
MiniMax does not publish a single hardware-independent generation-time figure. ComfyUI distributes BF16 and optimized/pruned variants. Generation time depends heavily on reference count, reference resolution, output length, diffusion steps, offloading and GPU.
About MiniMax H3 REF2VA + ComfyUI
What MiniMax H3 REF2VA + ComfyUI is built for, and where it falls short.
MiniMax H3 Ref2VA is the reference-conditioned branch of MiniMax H3. It jointly understands text, images, reference video and reference audio and generates synchronized video and stereo audio in a single multimodal generation system.
Ref2VA is not just image-to-video. The official interface can combine up to nine images, three video references and three standalone audio references, with no more than twelve reference files in total.
The generation backbone is a 33B dense single-stream H3 Omni Transformer. MiniMax states that roughly 13B parameters reside in AdaLN-related branches whose outputs can be precomputed and cached, reducing what must remain loaded during inference-only deployment.
For local workflows, H3-Base generates the 768p stage. MiniMax's full 2K pipeline regenerates that result with H3-Regenerate-2K rather than simply performing conventional spatial upscaling.
What people run it for
Character-consistent video
Generate new shots while using one or more images to establish a character's identity and visual appearance.
Voice-referenced scenes
Use audio references to guide voice timbre while jointly generating synchronized dialogue and video.
Motion and camera reference
Condition a new generation on reference video to communicate motion, framing or camera behavior.
Multi-reference creative generation
Combine separate subject, style, motion and audio references inside one multimodal generation request.
Strengths
- Can condition one generation on images, video clips and audio references together
- Designed for preserving subject identity, visual style, motion patterns, camera behavior and voice characteristics from references
- Generates video and synchronized stereo audio jointly instead of attaching a separately generated soundtrack afterward
- Supports multiple reference subjects and modalities in one request
- Native ComfyUI support includes a dedicated MiniMax H3 Reference to Video conditioning node
- Reference images can be processed at higher reference resolution for stronger fidelity when compute and memory permit
Limitations
- Reference consistency is a model capability rather than a guarantee; complex prompts and many competing references can still introduce visual or identity drift
- The open H3-Base workflow is the 768p generation stage; full 2K output requires the separate H3-Regenerate-2K workflow documented by MiniMax
- Reference tokens remain involved during sampling, so larger or more numerous references can materially increase compute and memory requirements
- The full BF16 model stack is extremely large; practical ComfyUI deployment commonly relies on pruned or quantized repackaged checkpoints
- MiniMax H3 uses a custom community licence with geographic and commercial conditions rather than a permissive Apache-style licence
Quickstart
Use ComfyUI's native MiniMax H3 Reference to Video workflow with the Ref2VA diffusion checkpoint, H3 Qwen3-VL text/vision encoder and both the video and audio VAEs.
1. Load minimax_h3_ref2va_pruned_int8_convrot.safetensors as the diffusion model.
2. Load the MiniMax-H3 Qwen3-VL-32B text/vision encoder.
3. Load minimax_h3_video_vae_fp16.safetensors and minimax_h3_audio_vae_fp32.safetensors.
4. Use the MiniMax H3 Reference to Video node.
5. Connect reference images, videos and/or audio.
6. Refer to them in the prompt as <Picture 1>, <Video 1>, <Audio 1>, etc.
7. Sample the joint audio-video latent and decode both streams before muxing the final video.ComfyUI's native node currently exposes up to nine image references, three video references and three standalone audio references.
MiniMax H3 REF2VA + ComfyUI specifications
Architecture and serving details for MiniMax H3 REF2VA + ComfyUI.
H3 architecture
- Generator
- H3 Omni Transformer
- Transformer parameters
- 33B dense
- Cacheable AdaLN-related parameters
- ~13B
- Text / vision encoder
- Qwen3-VL-32B, hidden state from layer 50
- Position encoding
- 3D Multimodal RoPE
References
- Reference images
- Up to 9
- Reference videos
- Up to 3
- Reference audio clips
- Up to 3
- Maximum total reference files
- 12
- Reference video length
- 2–15 seconds each; 15 seconds total
Output
- H3-Base output
- 768p-class, default short side 768 px
- Frame rate
- 24 FPS
- Duration
- 4–15 seconds
- Audio
- 32 kHz stereo
- 2K
- Separate H3-Regenerate-2K stage
ComfyUI
- Native node
- MiniMaxH3ReferenceToVideo
- Recommended local checkpoint family
- ref2va
- Video VAE
- minimax_h3_video_vae_fp16
- Audio VAE
- minimax_h3_audio_vae_fp32
MiniMax H3 REF2VA + ComfyUI frequently asked questions
Common questions about deploying MiniMax H3 REF2VA + ComfyUI on HexGrid Cloud.
Can Ref2VA use more than one reference image?
Yes. MiniMax officially supports up to nine reference images, and references can also include video and audio.
Does H3 generate audio natively?
Yes. H3 jointly predicts video and audio latents and outputs 32 kHz stereo audio synchronized with the generated video.
Can local H3 generate 2K directly?
The open H3-Base workflow produces the 768p generation stage. MiniMax's documented 2K pipeline uses the separate H3-Regenerate-2K stage.
Is MiniMax H3 Apache 2.0?
No. It uses the MiniMax H3 Community License Agreement, which includes geographic restrictions, acceptable-use requirements and additional commercial terms.
Are there commercial restrictions?
Yes. Among other terms, the licence excludes the US, EU, UK and South Korea from its standard territory and requires separate authorization when covered commercial products or services exceed USD 20 million in yearly revenue. Review the current licence itself before deployment.
Deploy MiniMax H3 REF2VA + ComfyUI today
$0.45 per hour on 1× RTX 4090, billed per minute, never shared.