MiniMax H3 TI2VA + ComfyUI
Deploy MiniMax H3 TI2VA + ComfyUI on a dedicated GPU from $0.45 per hour, billed per minute, on 1× RTX 4090. MiniMax H3's general text and frame-conditioned video checkpoint, generating synchronized video and stereo audio from text alone or optional first and last frame anchors.
HexGrid Cloud price
$0.45/hr
1× RTX 4090
Parameters
33B
dense
Modes
T2V + I2V
first / last frame
Duration
4–15s
24 FPS
Running MiniMax H3 TI2VA + ComfyUI on HexGrid Cloud
What you get when you deploy with us, beyond the hourly rate.
Single tenant by default
The model runs on a GPU that is yours for the life of the instance. No shared endpoint, no queue behind other tenants.
Per-minute billing
You pay for the GPU, not the token. Stop the instance and billing stops with it.
OpenAI-compatible endpoint
Point an existing SDK at your instance by changing the base URL. No rewrite required.
Your weights, your data
Prompts and outputs stay inside your instance. Nothing is logged, sampled or used for training.
MiniMax H3 TI2VA + ComfyUI GPU sizing and cost
What the model needs at each precision, and the cheapest HexGrid Cloud configuration that holds it.
| Precision | vRAM needed | Runs on | Price |
|---|---|---|---|
BF16 Recommended | 80 GB | 5× RTX A4000 | $1.00/hr |
FP8 | 40 GB | 1× RTX A6000 | $0.55/hr |
INT4 | 20 GB | 1× RTX 4090 | $0.45/hr |
Do not publish a generic seconds/video figure without fixing resolution, duration, sampling steps, checkpoint variant and GPU. ComfyUI provides optimized H3 repackagings suitable for offloaded local workflows.
About MiniMax H3 TI2VA + ComfyUI
What MiniMax H3 TI2VA + ComfyUI is built for, and where it falls short.
MiniMax H3 Base FL2VA is the open H3 checkpoint used for ordinary text-to-video-audio and image-anchored generation. MiniMax's official naming is FL2VA rather than TI2VA.
With no image it performs text-to-audio-video generation. A single image can establish either the first or last frame. Two images can constrain both ends of the shot.
The same H3 Omni Transformer jointly generates the video and stereo soundtrack, making speech, environmental sound effects and music part of the generation rather than a separate post-processing model.
ComfyUI exposes this family through its MiniMax H3 Image to Video conditioning node, whose first-frame and last-frame inputs are both optional.
What people run it for
Text-to-video with audio
Generate complete short-form audiovisual scenes directly from a text prompt.
Image-to-video
Use an image as the opening frame and generate natural motion that develops from it.
First-last frame transitions
Control both the beginning and ending appearance of a shot while the model generates the motion between them.
Last-frame-directed generation
Constrain where a generated scene should end while leaving the opening composition to the model.
Strengths
- One checkpoint handles text-only, first-frame, last-frame and first+last-frame generation
- Native synchronized stereo audio
- First and last frame constraints provide explicit shot-boundary control
- 24 FPS generation and up to 15-second clips
- Native ComfyUI implementation
- Uses the same H3 multimodal architecture as Ref2VA without the additional cost of arbitrary reference sets
Limitations
- This is officially FL2VA/T2VA, not a separately trained TI2VA checkpoint
- Single-image conditioning anchors a frame rather than functioning like Ref2VA's arbitrary subject/style reference system
- Local H3-Base generation is the 768p stage; MiniMax uses an additional H3-Regenerate-2K stage for 2K output
- 33B full weights plus the Qwen3-VL encoder and VAEs make the unquantized model stack very large
- The custom H3 community licence includes geographic and commercial restrictions
Quickstart
Use the H3 FL2VA checkpoint with ComfyUI's MiniMax H3 Image to Video node. Leave both frame inputs empty for T2VA, connect one for I2VA/L2VA, or connect both for FL2VA.
1. Load minimax_h3_fl2va_pruned_int8_convrot.safetensors.
2. Load the MiniMax-H3 Qwen3-VL-32B text/vision encoder.
3. Load the H3 video and audio VAEs.
4. Add MiniMax H3 Image to Video.
5. Text only: leave first_frame and last_frame empty.
6. First-frame I2V: connect first_frame.
7. Last-frame mode: connect last_frame.
8. First+last: connect both.
9. Sample and decode the joint video/audio latent.ComfyUI identifies this native conditioning node as MiniMaxH3ImageToVideo.
MiniMax H3 TI2VA + ComfyUI specifications
Architecture and serving details for MiniMax H3 TI2VA + ComfyUI.
Generation modes
- 0 images
- Text-to-Audio-Video
- 1 image
- First-frame or last-frame Video
- 2 images
- First-and-last-frame Video
Output
- Base resolution
- 768p-class
- Frame rate
- 24 FPS
- Duration
- 4–15 seconds
- Audio
- 32 kHz stereo
ComfyUI
- Conditioning node
- MiniMaxH3ImageToVideo
- Model family
- fl2va
- Default example canvas
- 1344 × 768
- Native FPS
- 24
MiniMax H3 TI2VA + ComfyUI frequently asked questions
Common questions about deploying MiniMax H3 TI2VA + ComfyUI on HexGrid Cloud.
Is TI2VA an official MiniMax checkpoint name?
No. MiniMax releases H3-Base-FL2VA, which covers text-to-audio-video and first/last-frame-conditioned generation. TI2VA can still be used as a catalogue label if the canonical mapping is stated.
Can it use a first and a last frame?
Yes. H3-Base-FL2VA supports zero, one or two images; two images define first and last frame conditions.
Does it create sound?
Yes. Video and 32 kHz stereo audio are generated jointly.
How is this different from Ref2VA?
FL2VA uses zero, one or two frame anchors. Ref2VA accepts arbitrary image, video and audio reference material for subject, style, motion or voice conditioning.
Deploy MiniMax H3 TI2VA + ComfyUI today
$0.45 per hour on 1× RTX 4090, billed per minute, never shared.