Wan2.2 A14B I2V + ComfyUI
Deploy Wan2.2 A14B I2V + ComfyUI on a dedicated GPU from $0.45 per hour, billed per minute, on 1× RTX 4090. Wan2.2's dual-expert image-to-video model, using a source image to anchor composition and appearance while separate high- and low-noise experts generate motion and refine detail.
HexGrid Cloud price
$0.45/hr
1× RTX 4090
Parameters
27B
~14B active
Input
Image + Text
image conditioning
Resolution
720p
also 480p
Running Wan2.2 A14B I2V + ComfyUI on HexGrid Cloud
What you get when you deploy with us, beyond the hourly rate.
Single tenant by default
The model runs on a GPU that is yours for the life of the instance. No shared endpoint, no queue behind other tenants.
Per-minute billing
You pay for the GPU, not the token. Stop the instance and billing stops with it.
OpenAI-compatible endpoint
Point an existing SDK at your instance by changing the base URL. No rewrite required.
Your weights, your data
Prompts and outputs stay inside your instance. Nothing is logged, sampled or used for training.
Wan2.2 A14B I2V + ComfyUI GPU sizing and cost
What the model needs at each precision, and the cheapest HexGrid Cloud configuration that holds it.
| Precision | vRAM needed | Runs on | Price |
|---|---|---|---|
BF16 Recommended | 65 GB | 5× RTX A4000 | $1.00/hr |
FP8 | 33 GB | 1× RTX A6000 | $0.55/hr |
INT4 | 17 GB | 1× RTX 4090 | $0.45/hr |
ComfyUI distributes separate high- and low-noise expert files. Actual VRAM varies significantly with precision, offloading, frame count and output resolution.
About Wan2.2 A14B I2V + ComfyUI
What Wan2.2 A14B I2V + ComfyUI is built for, and where it falls short.
Wan2.2 I2V-A14B is the image-to-video counterpart of Wan's dual-expert A14B text model. A source image provides the visual starting point and a text prompt guides motion, scene behavior and style.
As with T2V-A14B, the model uses separate high-noise and low-noise diffusion experts. The first establishes large-scale motion and structure while the second refines later denoising stages.
The full MoE contains approximately 27B parameters, although roughly 14B are active at any one denoising step.
The official implementation supports both 480p and 720p output and preserves the aspect ratio of the source image.
What people run it for
Animate product or character images
Turn a still composition into motion while retaining its main subjects and visual design.
Cinematic still-to-shot
Use a key visual as the opening frame and direct camera movement and scene action through a prompt.
First-last-frame transitions
Use ComfyUI's Wan2.2 first-last-frame workflow to specify both endpoints of a generated transition.
Stylized animation
Animate illustrations and stylized stills using the model's image-conditioned video generation.
Strengths
- Strong source-image conditioning for composition and subject appearance
- Dedicated I2V model rather than relying on a generalized T2V checkpoint
- Dual-expert design separates broad motion generation from detail refinement
- Supports both 480p and 720p generation
- Official first-frame and first-last-frame ComfyUI workflows
- Apache 2.0 licence
Limitations
- Image consistency is not mathematically guaranteed and can degrade under large transformations or difficult prompts
- The official reference command requires at least 80 GB GPU memory with documented single-GPU offloading settings
- 27B total weights make it substantially heavier than Wan2.2 TI2V-5B
- No synchronized native audio generation
- A14B should be described as the active parameter count, not total model size
Quickstart
Load the Wan2.2 I2V high- and low-noise diffusion checkpoints in ComfyUI, add an input image and encode the requested motion with the UMT5 text prompt.
1. Load wan2.2_i2v_high_noise_14B_fp16.safetensors.
2. Load wan2.2_i2v_low_noise_14B_fp16.safetensors.
3. Load the UMT5-XXL text encoder.
4. Load the Wan-compatible VAE.
5. Connect the source image to the I2V latent workflow.
6. Write the desired motion, camera and scene prompt.
7. Select frame count and target resolution.
8. Run the high-noise / low-noise expert sampling stages.ComfyUI also provides a first-last-frame workflow using the same I2V model family.
Wan2.2 A14B I2V + ComfyUI specifications
Architecture and serving details for Wan2.2 A14B I2V + ComfyUI.
Architecture
- Total parameters
- ~27B
- Active parameters
- ~14B
- Experts
- High-noise + low-noise
- Layers per expert
- 40
- Hidden dimension
- 5120
- Attention heads
- 40
Conditioning
- Input
- Image + optional text prompt
- Model input channels
- 36
- Text encoder dimension
- 4096
- Text length
- 512
Generation
- Resolution
- 480p / 720p
- Aspect ratio
- Follows source image
- Audio
- No native audio
Wan2.2 A14B I2V + ComfyUI frequently asked questions
Common questions about deploying Wan2.2 A14B I2V + ComfyUI on HexGrid Cloud.
Does Wan2.2 I2V preserve the source image exactly?
The source image strongly conditions composition and appearance, but generation can still alter fine details or identity under motion. Treat consistency as a capability rather than a strict guarantee.
Can it use a final frame?
ComfyUI includes a first-last-frame Wan2.2 workflow using the I2V model family, allowing both endpoints to be specified.
Why does the model have two large files?
Wan2.2 A14B uses separate high-noise and low-noise experts for different regions of the denoising process.
Is it really only 14B parameters?
No. Approximately 14B are active at a time, while the complete two-expert model contains roughly 27B parameters.
Deploy Wan2.2 A14B I2V + ComfyUI today
$0.45 per hour on 1× RTX 4090, billed per minute, never shared.