Wan2.2 5B TI2V + ComfyUI
Deploy Wan2.2 5B TI2V + ComfyUI on a dedicated GPU from $0.20 per hour, billed per minute, on 1× RTX A4000. A dense 5B Wan2.2 video generator that handles both text-to-video and image-to-video in one checkpoint, using a high-compression VAE to make 720p 24 FPS generation practical on consumer-class GPUs.
HexGrid Cloud price
$0.20/hr
1× RTX A4000
Parameters
5B
dense
Resolution
720p
1280 × 704 class
Frame rate
24 FPS
official
Running Wan2.2 5B TI2V + ComfyUI on HexGrid Cloud
What you get when you deploy with us, beyond the hourly rate.
Single tenant by default
The model runs on a GPU that is yours for the life of the instance. No shared endpoint, no queue behind other tenants.
Per-minute billing
You pay for the GPU, not the token. Stop the instance and billing stops with it.
OpenAI-compatible endpoint
Point an existing SDK at your instance by changing the base URL. No rewrite required.
Your weights, your data
Prompts and outputs stay inside your instance. Nothing is logged, sampled or used for training.
Wan2.2 5B TI2V + ComfyUI GPU sizing and cost
What the model needs at each precision, and the cheapest HexGrid Cloud configuration that holds it.
| Precision | vRAM needed | Runs on | Price |
|---|---|---|---|
BF16 Recommended | 12 GB | 1× RTX A4000 | $0.20/hr |
FP8 | 6 GB | 1× RTX A4000 | $0.20/hr |
INT4 | 3 GB | 1× RTX A4000 | $0.20/hr |
The <9-minute figure is Wan's own qualified reference claim, not a HexGrid benchmark. Do not surface it as guaranteed latency.
About Wan2.2 5B TI2V + ComfyUI
What Wan2.2 5B TI2V + ComfyUI is built for, and where it falls short.
Wan2.2 TI2V-5B is the compact unified model in the Wan2.2 release. A single dense 5B diffusion model supports both text-to-video and image-to-video rather than requiring separate task checkpoints.
Its efficiency comes partly from the Wan2.2 high-compression VAE. Wan reports a temporal-by-spatial compression ratio of 4×16×16, increasing to an effective 4×32×32 after patchification.
The official model generates 720p-class video at 24 FPS and can run on an RTX 4090-class 24 GB GPU using CPU text-encoder placement, dtype conversion and model offloading.
ComfyUI's native Wan2.2 workflow is significantly simpler than A14B because it loads one 5B diffusion model rather than separate high- and low-noise experts.
What people run it for
Consumer-GPU video generation
Run 720p text- or image-conditioned video generation on a 24GB-class GPU with offloading.
Rapid creative prototyping
Use one model for both prompt-only generations and animation of source images.
Local ComfyUI workflows
Use a straightforward one-model Wan2.2 graph instead of coordinating two A14B expert checkpoints.
Private image animation
Animate sensitive or proprietary source imagery entirely on dedicated infrastructure.
Strengths
- Single checkpoint handles both text-to-video and image-to-video
- Much smaller deployment footprint than Wan2.2 A14B
- Official 24GB single-GPU reference configuration
- 720p 24 FPS output
- High-compression VAE reduces video-latent compute
- Simple native ComfyUI workflow
- Apache 2.0
Limitations
- Lower overall model capacity than the A14B dual-expert checkpoints
- Wan's consumer-GPU configuration relies on model offloading and moving the T5 encoder to CPU
- Official 720p dimensions for TI2V use 1280×704 or 704×1280 rather than exact 1280×720
- Wan reports under nine minutes for a five-second 720p clip on a consumer GPU without specialized optimization; this remains much slower than real-time generation
- No native synchronized audio
Quickstart
ComfyUI uses one 5B diffusion checkpoint, an FP8 UMT5-XXL text encoder and the Wan2.2 VAE. Connect an image only when image-to-video generation is required.
1. Load wan2.2_ti2v_5B_fp16.safetensors.
2. Load umt5_xxl_fp8_e4m3fn_scaled.safetensors.
3. Load wan2.2_vae.safetensors.
4. For T2V, enter a prompt and leave image conditioning disabled.
5. For I2V, connect a source image to Wan22ImageToVideoLatent.
6. Set target dimensions and frame length.
7. Run the Wan2.2 sampling workflow and encode the resulting frames as video.ComfyUI states the 5B workflow should fit well on 8GB VRAM with its native offloading, while Wan's original reference CLI documents 24GB as the minimum for its own single-GPU recipe. These are different runtime configurations and should not be presented as the same requirement.
Wan2.2 5B TI2V + ComfyUI specifications
Architecture and serving details for Wan2.2 5B TI2V + ComfyUI.
Architecture
- Type
- Dense video diffusion Transformer
- Parameters
- 5B
- Layers
- 30
- Hidden dimension
- 3072
- FFN dimension
- 14336
- Attention heads
- 24
Video compression
- VAE compression
- 4 × 16 × 16
- With patchification
- 4 × 32 × 32
- Latent channels
- 48
Generation
- Tasks
- Text-to-video + image-to-video
- Resolution
- 720p class
- Official dimensions
- 1280×704 / 704×1280
- Frame rate
- 24 FPS
- Audio
- No
ComfyUI
- Diffusion model
- wan2.2_ti2v_5B_fp16
- Text encoder
- UMT5-XXL FP8
- VAE
- Wan2.2 VAE
- Native offloading
- Supported
Wan2.2 5B TI2V + ComfyUI frequently asked questions
Common questions about deploying Wan2.2 5B TI2V + ComfyUI on HexGrid Cloud.
Can the same 5B model do both T2V and I2V?
Yes. Wan explicitly describes TI2V-5B as a unified model supporting both text-to-video and image-to-video.
Can it run on an RTX 4090?
Wan's reference implementation supports a 24GB RTX 4090-class GPU using model offloading, dtype conversion and CPU placement of the T5 encoder.
Why does ComfyUI sometimes claim only 8GB is needed?
ComfyUI's documentation says its native offloading workflow should fit well on 8GB VRAM. That is a ComfyUI-specific memory-management configuration and should be distinguished from Wan's own 24GB reference CLI requirement.
How fast is generation?
Wan reports that without special optimization, a five-second 720p video can be generated in under nine minutes on a single consumer-grade GPU. Actual speed varies substantially by GPU and runtime.
Does it generate audio?
No. Wan2.2 TI2V-5B generates video frames but does not jointly generate synchronized audio.
Deploy Wan2.2 5B TI2V + ComfyUI today
$0.20 per hour on 1× RTX A4000, billed per minute, never shared.