Back to all models

MiniMax H3 REF2VA + ComfyUI

Deploy MiniMax H3 REF2VA + ComfyUI on a dedicated GPU from $0.45 per hour, billed per minute, on 1× RTX 4090. MiniMax H3's reference-conditioned audio-video checkpoint, built to generate new scenes from text plus image, video and audio references while preserving subject, style, motion and voice cues.

Text33B paramsMiniMax H3 Community License AgreementDeploys in ~5 minSecure Cloud
MiniMaxAI/MiniMax-H3

HexGrid Cloud price

$0.45/hr

1× RTX 4090

Parameters

33B

dense Omni Transformer

Output

768p

local H3-Base stage

Duration

4–15s

24 FPS

Running MiniMax H3 REF2VA + ComfyUI on HexGrid Cloud

What you get when you deploy with us, beyond the hourly rate.

Single tenant by default

The model runs on a GPU that is yours for the life of the instance. No shared endpoint, no queue behind other tenants.

Per-minute billing

You pay for the GPU, not the token. Stop the instance and billing stops with it.

OpenAI-compatible endpoint

Point an existing SDK at your instance by changing the base URL. No rewrite required.

Your weights, your data

Prompts and outputs stay inside your instance. Nothing is logged, sampled or used for training.

MiniMax H3 REF2VA + ComfyUI GPU sizing and cost

What the model needs at each precision, and the cheapest HexGrid Cloud configuration that holds it.

MiniMax H3 REF2VA + ComfyUI vRAM requirement and hourly cost by precision
PrecisionvRAM neededRuns onPrice
BF16
Recommended
80 GB5× RTX A4000$1.00/hr
FP8
40 GB1× RTX A6000$0.55/hr
INT4
20 GB1× RTX 4090$0.45/hr

MiniMax does not publish a single hardware-independent generation-time figure. ComfyUI distributes BF16 and optimized/pruned variants. Generation time depends heavily on reference count, reference resolution, output length, diffusion steps, offloading and GPU.

About MiniMax H3 REF2VA + ComfyUI

What MiniMax H3 REF2VA + ComfyUI is built for, and where it falls short.

MiniMax H3 Ref2VA is the reference-conditioned branch of MiniMax H3. It jointly understands text, images, reference video and reference audio and generates synchronized video and stereo audio in a single multimodal generation system.

Ref2VA is not just image-to-video. The official interface can combine up to nine images, three video references and three standalone audio references, with no more than twelve reference files in total.

The generation backbone is a 33B dense single-stream H3 Omni Transformer. MiniMax states that roughly 13B parameters reside in AdaLN-related branches whose outputs can be precomputed and cached, reducing what must remain loaded during inference-only deployment.

For local workflows, H3-Base generates the 768p stage. MiniMax's full 2K pipeline regenerates that result with H3-Regenerate-2K rather than simply performing conventional spatial upscaling.

What people run it for

Character-consistent video

Generate new shots while using one or more images to establish a character's identity and visual appearance.

Voice-referenced scenes

Use audio references to guide voice timbre while jointly generating synchronized dialogue and video.

Motion and camera reference

Condition a new generation on reference video to communicate motion, framing or camera behavior.

Multi-reference creative generation

Combine separate subject, style, motion and audio references inside one multimodal generation request.

Strengths

  • Can condition one generation on images, video clips and audio references together
  • Designed for preserving subject identity, visual style, motion patterns, camera behavior and voice characteristics from references
  • Generates video and synchronized stereo audio jointly instead of attaching a separately generated soundtrack afterward
  • Supports multiple reference subjects and modalities in one request
  • Native ComfyUI support includes a dedicated MiniMax H3 Reference to Video conditioning node
  • Reference images can be processed at higher reference resolution for stronger fidelity when compute and memory permit

Limitations

  • Reference consistency is a model capability rather than a guarantee; complex prompts and many competing references can still introduce visual or identity drift
  • The open H3-Base workflow is the 768p generation stage; full 2K output requires the separate H3-Regenerate-2K workflow documented by MiniMax
  • Reference tokens remain involved during sampling, so larger or more numerous references can materially increase compute and memory requirements
  • The full BF16 model stack is extremely large; practical ComfyUI deployment commonly relies on pruned or quantized repackaged checkpoints
  • MiniMax H3 uses a custom community licence with geographic and commercial conditions rather than a permissive Apache-style licence

Quickstart

Use ComfyUI's native MiniMax H3 Reference to Video workflow with the Ref2VA diffusion checkpoint, H3 Qwen3-VL text/vision encoder and both the video and audio VAEs.

comfyuiComfyUI workflow
1. Load minimax_h3_ref2va_pruned_int8_convrot.safetensors as the diffusion model.
2. Load the MiniMax-H3 Qwen3-VL-32B text/vision encoder.
3. Load minimax_h3_video_vae_fp16.safetensors and minimax_h3_audio_vae_fp32.safetensors.
4. Use the MiniMax H3 Reference to Video node.
5. Connect reference images, videos and/or audio.
6. Refer to them in the prompt as <Picture 1>, <Video 1>, <Audio 1>, etc.
7. Sample the joint audio-video latent and decode both streams before muxing the final video.

ComfyUI's native node currently exposes up to nine image references, three video references and three standalone audio references.

MiniMax H3 REF2VA + ComfyUI specifications

Architecture and serving details for MiniMax H3 REF2VA + ComfyUI.

PublisherMiniMaxParameters33BLicenceMiniMax H3 Community License AgreementReleasedJuly 31, 2026

H3 architecture

Generator
H3 Omni Transformer
Transformer parameters
33B dense
Cacheable AdaLN-related parameters
~13B
Text / vision encoder
Qwen3-VL-32B, hidden state from layer 50
Position encoding
3D Multimodal RoPE

References

Reference images
Up to 9
Reference videos
Up to 3
Reference audio clips
Up to 3
Maximum total reference files
12
Reference video length
2–15 seconds each; 15 seconds total

Output

H3-Base output
768p-class, default short side 768 px
Frame rate
24 FPS
Duration
4–15 seconds
Audio
32 kHz stereo
2K
Separate H3-Regenerate-2K stage

ComfyUI

Native node
MiniMaxH3ReferenceToVideo
Recommended local checkpoint family
ref2va
Video VAE
minimax_h3_video_vae_fp16
Audio VAE
minimax_h3_audio_vae_fp32

MiniMax H3 REF2VA + ComfyUI frequently asked questions

Common questions about deploying MiniMax H3 REF2VA + ComfyUI on HexGrid Cloud.

Can Ref2VA use more than one reference image?

Yes. MiniMax officially supports up to nine reference images, and references can also include video and audio.

Does H3 generate audio natively?

Yes. H3 jointly predicts video and audio latents and outputs 32 kHz stereo audio synchronized with the generated video.

Can local H3 generate 2K directly?

The open H3-Base workflow produces the 768p generation stage. MiniMax's documented 2K pipeline uses the separate H3-Regenerate-2K stage.

Is MiniMax H3 Apache 2.0?

No. It uses the MiniMax H3 Community License Agreement, which includes geographic restrictions, acceptable-use requirements and additional commercial terms.

Are there commercial restrictions?

Yes. Among other terms, the licence excludes the US, EU, UK and South Korea from its standard territory and requires separate authorization when covered commercial products or services exceed USD 20 million in yearly revenue. Review the current licence itself before deployment.

Deploy MiniMax H3 REF2VA + ComfyUI today

$0.45 per hour on 1× RTX 4090, billed per minute, never shared.