Back to all models

Wan2.2 A14B I2V + ComfyUI

Deploy Wan2.2 A14B I2V + ComfyUI on a dedicated GPU from $0.45 per hour, billed per minute, on 1× RTX 4090. Wan2.2's dual-expert image-to-video model, using a source image to anchor composition and appearance while separate high- and low-noise experts generate motion and refine detail.

Text27B paramsApache 2.0Deploys in ~5 minSecure Cloud
Wan-AI/Wan2.2-I2V-A14B

HexGrid Cloud price

$0.45/hr

1× RTX 4090

Parameters

27B

~14B active

Input

Image + Text

image conditioning

Resolution

720p

also 480p

Running Wan2.2 A14B I2V + ComfyUI on HexGrid Cloud

What you get when you deploy with us, beyond the hourly rate.

Single tenant by default

The model runs on a GPU that is yours for the life of the instance. No shared endpoint, no queue behind other tenants.

Per-minute billing

You pay for the GPU, not the token. Stop the instance and billing stops with it.

OpenAI-compatible endpoint

Point an existing SDK at your instance by changing the base URL. No rewrite required.

Your weights, your data

Prompts and outputs stay inside your instance. Nothing is logged, sampled or used for training.

Wan2.2 A14B I2V + ComfyUI GPU sizing and cost

What the model needs at each precision, and the cheapest HexGrid Cloud configuration that holds it.

Wan2.2 A14B I2V + ComfyUI vRAM requirement and hourly cost by precision
PrecisionvRAM neededRuns onPrice
BF16
Recommended
65 GB5× RTX A4000$1.00/hr
FP8
33 GB1× RTX A6000$0.55/hr
INT4
17 GB1× RTX 4090$0.45/hr

ComfyUI distributes separate high- and low-noise expert files. Actual VRAM varies significantly with precision, offloading, frame count and output resolution.

About Wan2.2 A14B I2V + ComfyUI

What Wan2.2 A14B I2V + ComfyUI is built for, and where it falls short.

Wan2.2 I2V-A14B is the image-to-video counterpart of Wan's dual-expert A14B text model. A source image provides the visual starting point and a text prompt guides motion, scene behavior and style.

As with T2V-A14B, the model uses separate high-noise and low-noise diffusion experts. The first establishes large-scale motion and structure while the second refines later denoising stages.

The full MoE contains approximately 27B parameters, although roughly 14B are active at any one denoising step.

The official implementation supports both 480p and 720p output and preserves the aspect ratio of the source image.

What people run it for

Animate product or character images

Turn a still composition into motion while retaining its main subjects and visual design.

Cinematic still-to-shot

Use a key visual as the opening frame and direct camera movement and scene action through a prompt.

First-last-frame transitions

Use ComfyUI's Wan2.2 first-last-frame workflow to specify both endpoints of a generated transition.

Stylized animation

Animate illustrations and stylized stills using the model's image-conditioned video generation.

Strengths

  • Strong source-image conditioning for composition and subject appearance
  • Dedicated I2V model rather than relying on a generalized T2V checkpoint
  • Dual-expert design separates broad motion generation from detail refinement
  • Supports both 480p and 720p generation
  • Official first-frame and first-last-frame ComfyUI workflows
  • Apache 2.0 licence

Limitations

  • Image consistency is not mathematically guaranteed and can degrade under large transformations or difficult prompts
  • The official reference command requires at least 80 GB GPU memory with documented single-GPU offloading settings
  • 27B total weights make it substantially heavier than Wan2.2 TI2V-5B
  • No synchronized native audio generation
  • A14B should be described as the active parameter count, not total model size

Quickstart

Load the Wan2.2 I2V high- and low-noise diffusion checkpoints in ComfyUI, add an input image and encode the requested motion with the UMT5 text prompt.

comfyuiComfyUI workflow
1. Load wan2.2_i2v_high_noise_14B_fp16.safetensors.
2. Load wan2.2_i2v_low_noise_14B_fp16.safetensors.
3. Load the UMT5-XXL text encoder.
4. Load the Wan-compatible VAE.
5. Connect the source image to the I2V latent workflow.
6. Write the desired motion, camera and scene prompt.
7. Select frame count and target resolution.
8. Run the high-noise / low-noise expert sampling stages.

ComfyUI also provides a first-last-frame workflow using the same I2V model family.

Wan2.2 A14B I2V + ComfyUI specifications

Architecture and serving details for Wan2.2 A14B I2V + ComfyUI.

PublisherAlibaba WanParameters27BLicenceApache 2.0ReleasedJuly 28, 2025

Architecture

Total parameters
~27B
Active parameters
~14B
Experts
High-noise + low-noise
Layers per expert
40
Hidden dimension
5120
Attention heads
40

Conditioning

Input
Image + optional text prompt
Model input channels
36
Text encoder dimension
4096
Text length
512

Generation

Resolution
480p / 720p
Aspect ratio
Follows source image
Audio
No native audio

Wan2.2 A14B I2V + ComfyUI frequently asked questions

Common questions about deploying Wan2.2 A14B I2V + ComfyUI on HexGrid Cloud.

Does Wan2.2 I2V preserve the source image exactly?

The source image strongly conditions composition and appearance, but generation can still alter fine details or identity under motion. Treat consistency as a capability rather than a strict guarantee.

Can it use a final frame?

ComfyUI includes a first-last-frame Wan2.2 workflow using the I2V model family, allowing both endpoints to be specified.

Why does the model have two large files?

Wan2.2 A14B uses separate high-noise and low-noise experts for different regions of the denoising process.

Is it really only 14B parameters?

No. Approximately 14B are active at a time, while the complete two-expert model contains roughly 27B parameters.

Deploy Wan2.2 A14B I2V + ComfyUI today

$0.45 per hour on 1× RTX 4090, billed per minute, never shared.