Back to all models

MiniMax H3 TI2VA + ComfyUI

Deploy MiniMax H3 TI2VA + ComfyUI on a dedicated GPU from $0.45 per hour, billed per minute, on 1× RTX 4090. MiniMax H3's general text and frame-conditioned video checkpoint, generating synchronized video and stereo audio from text alone or optional first and last frame anchors.

Text33B paramsMiniMax H3 Community License AgreementDeploys in ~5 minSecure Cloud
MiniMaxAI/MiniMax-H3

HexGrid Cloud price

$0.45/hr

1× RTX 4090

Parameters

33B

dense

Modes

T2V + I2V

first / last frame

Duration

4–15s

24 FPS

Running MiniMax H3 TI2VA + ComfyUI on HexGrid Cloud

What you get when you deploy with us, beyond the hourly rate.

Single tenant by default

The model runs on a GPU that is yours for the life of the instance. No shared endpoint, no queue behind other tenants.

Per-minute billing

You pay for the GPU, not the token. Stop the instance and billing stops with it.

OpenAI-compatible endpoint

Point an existing SDK at your instance by changing the base URL. No rewrite required.

Your weights, your data

Prompts and outputs stay inside your instance. Nothing is logged, sampled or used for training.

MiniMax H3 TI2VA + ComfyUI GPU sizing and cost

What the model needs at each precision, and the cheapest HexGrid Cloud configuration that holds it.

MiniMax H3 TI2VA + ComfyUI vRAM requirement and hourly cost by precision
PrecisionvRAM neededRuns onPrice
BF16
Recommended
80 GB5× RTX A4000$1.00/hr
FP8
40 GB1× RTX A6000$0.55/hr
INT4
20 GB1× RTX 4090$0.45/hr

Do not publish a generic seconds/video figure without fixing resolution, duration, sampling steps, checkpoint variant and GPU. ComfyUI provides optimized H3 repackagings suitable for offloaded local workflows.

About MiniMax H3 TI2VA + ComfyUI

What MiniMax H3 TI2VA + ComfyUI is built for, and where it falls short.

MiniMax H3 Base FL2VA is the open H3 checkpoint used for ordinary text-to-video-audio and image-anchored generation. MiniMax's official naming is FL2VA rather than TI2VA.

With no image it performs text-to-audio-video generation. A single image can establish either the first or last frame. Two images can constrain both ends of the shot.

The same H3 Omni Transformer jointly generates the video and stereo soundtrack, making speech, environmental sound effects and music part of the generation rather than a separate post-processing model.

ComfyUI exposes this family through its MiniMax H3 Image to Video conditioning node, whose first-frame and last-frame inputs are both optional.

What people run it for

Text-to-video with audio

Generate complete short-form audiovisual scenes directly from a text prompt.

Image-to-video

Use an image as the opening frame and generate natural motion that develops from it.

First-last frame transitions

Control both the beginning and ending appearance of a shot while the model generates the motion between them.

Last-frame-directed generation

Constrain where a generated scene should end while leaving the opening composition to the model.

Strengths

  • One checkpoint handles text-only, first-frame, last-frame and first+last-frame generation
  • Native synchronized stereo audio
  • First and last frame constraints provide explicit shot-boundary control
  • 24 FPS generation and up to 15-second clips
  • Native ComfyUI implementation
  • Uses the same H3 multimodal architecture as Ref2VA without the additional cost of arbitrary reference sets

Limitations

  • This is officially FL2VA/T2VA, not a separately trained TI2VA checkpoint
  • Single-image conditioning anchors a frame rather than functioning like Ref2VA's arbitrary subject/style reference system
  • Local H3-Base generation is the 768p stage; MiniMax uses an additional H3-Regenerate-2K stage for 2K output
  • 33B full weights plus the Qwen3-VL encoder and VAEs make the unquantized model stack very large
  • The custom H3 community licence includes geographic and commercial restrictions

Quickstart

Use the H3 FL2VA checkpoint with ComfyUI's MiniMax H3 Image to Video node. Leave both frame inputs empty for T2VA, connect one for I2VA/L2VA, or connect both for FL2VA.

comfyuiComfyUI workflow
1. Load minimax_h3_fl2va_pruned_int8_convrot.safetensors.
2. Load the MiniMax-H3 Qwen3-VL-32B text/vision encoder.
3. Load the H3 video and audio VAEs.
4. Add MiniMax H3 Image to Video.
5. Text only: leave first_frame and last_frame empty.
6. First-frame I2V: connect first_frame.
7. Last-frame mode: connect last_frame.
8. First+last: connect both.
9. Sample and decode the joint video/audio latent.

ComfyUI identifies this native conditioning node as MiniMaxH3ImageToVideo.

MiniMax H3 TI2VA + ComfyUI specifications

Architecture and serving details for MiniMax H3 TI2VA + ComfyUI.

PublisherMiniMaxParameters33BLicenceMiniMax H3 Community License AgreementReleasedJuly 31, 2026

Generation modes

0 images
Text-to-Audio-Video
1 image
First-frame or last-frame Video
2 images
First-and-last-frame Video

Output

Base resolution
768p-class
Frame rate
24 FPS
Duration
4–15 seconds
Audio
32 kHz stereo

ComfyUI

Conditioning node
MiniMaxH3ImageToVideo
Model family
fl2va
Default example canvas
1344 × 768
Native FPS
24

MiniMax H3 TI2VA + ComfyUI frequently asked questions

Common questions about deploying MiniMax H3 TI2VA + ComfyUI on HexGrid Cloud.

Is TI2VA an official MiniMax checkpoint name?

No. MiniMax releases H3-Base-FL2VA, which covers text-to-audio-video and first/last-frame-conditioned generation. TI2VA can still be used as a catalogue label if the canonical mapping is stated.

Can it use a first and a last frame?

Yes. H3-Base-FL2VA supports zero, one or two images; two images define first and last frame conditions.

Does it create sound?

Yes. Video and 32 kHz stereo audio are generated jointly.

How is this different from Ref2VA?

FL2VA uses zero, one or two frame anchors. Ref2VA accepts arbitrary image, video and audio reference material for subject, style, motion or voice conditioning.

Deploy MiniMax H3 TI2VA + ComfyUI today

$0.45 per hour on 1× RTX 4090, billed per minute, never shared.