Back to all models

Wan2.2 5B TI2V + ComfyUI

Deploy Wan2.2 5B TI2V + ComfyUI on a dedicated GPU from $0.20 per hour, billed per minute, on 1× RTX A4000. A dense 5B Wan2.2 video generator that handles both text-to-video and image-to-video in one checkpoint, using a high-compression VAE to make 720p 24 FPS generation practical on consumer-class GPUs.

Text5B paramsApache 2.0Deploys in ~5 minSecure Cloud
Wan-AI/Wan2.2-TI2V-5B

HexGrid Cloud price

$0.20/hr

1× RTX A4000

Parameters

5B

dense

Resolution

720p

1280 × 704 class

Frame rate

24 FPS

official

Running Wan2.2 5B TI2V + ComfyUI on HexGrid Cloud

What you get when you deploy with us, beyond the hourly rate.

Single tenant by default

The model runs on a GPU that is yours for the life of the instance. No shared endpoint, no queue behind other tenants.

Per-minute billing

You pay for the GPU, not the token. Stop the instance and billing stops with it.

OpenAI-compatible endpoint

Point an existing SDK at your instance by changing the base URL. No rewrite required.

Your weights, your data

Prompts and outputs stay inside your instance. Nothing is logged, sampled or used for training.

Wan2.2 5B TI2V + ComfyUI GPU sizing and cost

What the model needs at each precision, and the cheapest HexGrid Cloud configuration that holds it.

Wan2.2 5B TI2V + ComfyUI vRAM requirement and hourly cost by precision
PrecisionvRAM neededRuns onPrice
BF16
Recommended
12 GB1× RTX A4000$0.20/hr
FP8
6 GB1× RTX A4000$0.20/hr
INT4
3 GB1× RTX A4000$0.20/hr

The <9-minute figure is Wan's own qualified reference claim, not a HexGrid benchmark. Do not surface it as guaranteed latency.

About Wan2.2 5B TI2V + ComfyUI

What Wan2.2 5B TI2V + ComfyUI is built for, and where it falls short.

Wan2.2 TI2V-5B is the compact unified model in the Wan2.2 release. A single dense 5B diffusion model supports both text-to-video and image-to-video rather than requiring separate task checkpoints.

Its efficiency comes partly from the Wan2.2 high-compression VAE. Wan reports a temporal-by-spatial compression ratio of 4×16×16, increasing to an effective 4×32×32 after patchification.

The official model generates 720p-class video at 24 FPS and can run on an RTX 4090-class 24 GB GPU using CPU text-encoder placement, dtype conversion and model offloading.

ComfyUI's native Wan2.2 workflow is significantly simpler than A14B because it loads one 5B diffusion model rather than separate high- and low-noise experts.

What people run it for

Consumer-GPU video generation

Run 720p text- or image-conditioned video generation on a 24GB-class GPU with offloading.

Rapid creative prototyping

Use one model for both prompt-only generations and animation of source images.

Local ComfyUI workflows

Use a straightforward one-model Wan2.2 graph instead of coordinating two A14B expert checkpoints.

Private image animation

Animate sensitive or proprietary source imagery entirely on dedicated infrastructure.

Strengths

  • Single checkpoint handles both text-to-video and image-to-video
  • Much smaller deployment footprint than Wan2.2 A14B
  • Official 24GB single-GPU reference configuration
  • 720p 24 FPS output
  • High-compression VAE reduces video-latent compute
  • Simple native ComfyUI workflow
  • Apache 2.0

Limitations

  • Lower overall model capacity than the A14B dual-expert checkpoints
  • Wan's consumer-GPU configuration relies on model offloading and moving the T5 encoder to CPU
  • Official 720p dimensions for TI2V use 1280×704 or 704×1280 rather than exact 1280×720
  • Wan reports under nine minutes for a five-second 720p clip on a consumer GPU without specialized optimization; this remains much slower than real-time generation
  • No native synchronized audio

Quickstart

ComfyUI uses one 5B diffusion checkpoint, an FP8 UMT5-XXL text encoder and the Wan2.2 VAE. Connect an image only when image-to-video generation is required.

comfyuiComfyUI workflow
1. Load wan2.2_ti2v_5B_fp16.safetensors.
2. Load umt5_xxl_fp8_e4m3fn_scaled.safetensors.
3. Load wan2.2_vae.safetensors.
4. For T2V, enter a prompt and leave image conditioning disabled.
5. For I2V, connect a source image to Wan22ImageToVideoLatent.
6. Set target dimensions and frame length.
7. Run the Wan2.2 sampling workflow and encode the resulting frames as video.

ComfyUI states the 5B workflow should fit well on 8GB VRAM with its native offloading, while Wan's original reference CLI documents 24GB as the minimum for its own single-GPU recipe. These are different runtime configurations and should not be presented as the same requirement.

Wan2.2 5B TI2V + ComfyUI specifications

Architecture and serving details for Wan2.2 5B TI2V + ComfyUI.

PublisherAlibaba WanParameters5BLicenceApache 2.0ReleasedJuly 28, 2025

Architecture

Type
Dense video diffusion Transformer
Parameters
5B
Layers
30
Hidden dimension
3072
FFN dimension
14336
Attention heads
24

Video compression

VAE compression
4 × 16 × 16
With patchification
4 × 32 × 32
Latent channels
48

Generation

Tasks
Text-to-video + image-to-video
Resolution
720p class
Official dimensions
1280×704 / 704×1280
Frame rate
24 FPS
Audio
No

ComfyUI

Diffusion model
wan2.2_ti2v_5B_fp16
Text encoder
UMT5-XXL FP8
VAE
Wan2.2 VAE
Native offloading
Supported

Wan2.2 5B TI2V + ComfyUI frequently asked questions

Common questions about deploying Wan2.2 5B TI2V + ComfyUI on HexGrid Cloud.

Can the same 5B model do both T2V and I2V?

Yes. Wan explicitly describes TI2V-5B as a unified model supporting both text-to-video and image-to-video.

Can it run on an RTX 4090?

Wan's reference implementation supports a 24GB RTX 4090-class GPU using model offloading, dtype conversion and CPU placement of the T5 encoder.

Why does ComfyUI sometimes claim only 8GB is needed?

ComfyUI's documentation says its native offloading workflow should fit well on 8GB VRAM. That is a ComfyUI-specific memory-management configuration and should be distinguished from Wan's own 24GB reference CLI requirement.

How fast is generation?

Wan reports that without special optimization, a five-second 720p video can be generated in under nine minutes on a single consumer-grade GPU. Actual speed varies substantially by GPU and runtime.

Does it generate audio?

No. Wan2.2 TI2V-5B generates video frames but does not jointly generate synchronized audio.

Deploy Wan2.2 5B TI2V + ComfyUI today

$0.20 per hour on 1× RTX A4000, billed per minute, never shared.