Wan2.2 A14B T2V + ComfyUI
Deploy Wan2.2 A14B T2V + ComfyUI on a dedicated GPU from $0.45 per hour, billed per minute, on 1× RTX 4090. Wan2.2's 27B-total, 14B-active text-to-video MoE, using separate high-noise and low-noise experts to handle scene layout and fine visual detail across the diffusion trajectory.
HexGrid Cloud price
$0.45/hr
1× RTX 4090
Parameters
27B
~14B active
Experts
2
high + low noise
Resolution
720p
also 480p
Running Wan2.2 A14B T2V + ComfyUI on HexGrid Cloud
What you get when you deploy with us, beyond the hourly rate.
Single tenant by default
The model runs on a GPU that is yours for the life of the instance. No shared endpoint, no queue behind other tenants.
Per-minute billing
You pay for the GPU, not the token. Stop the instance and billing stops with it.
OpenAI-compatible endpoint
Point an existing SDK at your instance by changing the base URL. No rewrite required.
Your weights, your data
Prompts and outputs stay inside your instance. Nothing is logged, sampled or used for training.
Wan2.2 A14B T2V + ComfyUI GPU sizing and cost
What the model needs at each precision, and the cheapest HexGrid Cloud configuration that holds it.
| Precision | vRAM needed | Runs on | Price |
|---|---|---|---|
BF16 Recommended | 65 GB | 5× RTX A4000 | $1.00/hr |
FP8 | 33 GB | 1× RTX A6000 | $0.55/hr |
INT4 | 17 GB | 1× RTX 4090 | $0.45/hr |
Generation time depends on resolution, frame count, sampling configuration and GPU. ComfyUI commonly uses separately quantized high-noise and low-noise expert files.
About Wan2.2 A14B T2V + ComfyUI
What Wan2.2 A14B T2V + ComfyUI is built for, and where it falls short.
Wan2.2 T2V-A14B is Alibaba Wan's dedicated text-to-video mixture-of-experts checkpoint. Its two approximately 14B diffusion experts specialize in different portions of the denoising trajectory.
The high-noise expert handles early denoising stages where global structure, composition and overall motion are established. The low-noise expert takes over later to refine visual detail.
Together the two experts contain approximately 27B parameters, while only about 14B are active at each denoising step.
The official model supports both 480p and 720p generation and has native support in ComfyUI with separate high-noise and low-noise diffusion-model files.
What people run it for
Cinematic text-to-video
Generate 480p or 720p shots from detailed natural-language scene and camera descriptions.
Motion-heavy scenes
Create clips where complex movement and camera dynamics are central to the prompt.
Creative previsualization
Prototype film, advertising and visual-concept shots directly from written treatments.
Private video generation
Run an Apache-licensed high-capacity text-to-video model on dedicated GPU infrastructure.
Strengths
- Specialized high-noise and low-noise experts split global composition and fine-detail generation
- Strong emphasis on cinematic composition, lighting, color and camera aesthetics
- Training data increased substantially over Wan2.1, including 83.2% more video data according to Wan
- Dedicated text-to-video checkpoint rather than a smaller unified compromise model
- 720p and 480p output support
- Apache 2.0 open-weight licence
- Official native ComfyUI workflow
Limitations
- 27B total weights make deployment substantially heavier than the 5B TI2V variant
- The official single-GPU reference implementation requires at least 80 GB VRAM even with its documented offload/conversion settings
- No native audio generation; generated output is visual video
- A14B refers to active compute rather than total model size and can easily be misrepresented in catalogue metadata
- Text-to-video only; image conditioning requires the I2V checkpoint
Quickstart
ComfyUI's official native workflow loads the high-noise and low-noise experts separately, plus the Wan-compatible VAE and UMT5 text encoder.
1. Load wan2.2_t2v_high_noise_14B_fp8_scaled.safetensors.
2. Load wan2.2_t2v_low_noise_14B_fp8_scaled.safetensors.
3. Load umt5_xxl_fp8_e4m3fn_scaled.safetensors.
4. Load wan_2.1_vae.safetensors.
5. Set width, height and frame count in the video latent node.
6. Enter positive and negative prompts.
7. Run the dual-expert Wan2.2 sampling workflow.The ComfyUI packaging names each expert 14B because each individual expert is approximately 14B; the complete A14B model has approximately 27B total parameters.
Wan2.2 A14B T2V + ComfyUI specifications
Architecture and serving details for Wan2.2 A14B T2V + ComfyUI.
MoE architecture
- Total parameters
- ~27B
- Active parameters
- ~14B per denoising step
- Experts
- 2
- Early denoising
- High-noise expert
- Late denoising
- Low-noise expert
Each expert
- Transformer layers
- 40
- Hidden dimension
- 5120
- FFN dimension
- 13824
- Attention heads
- 40
- Text length
- 512
Generation
- Task
- Text-to-video
- Supported resolution
- 480p / 720p
- Official example
- 1280 × 720
- Audio generation
- No
ComfyUI
- High-noise model
- Separate diffusion checkpoint
- Low-noise model
- Separate diffusion checkpoint
- Text encoder
- UMT5-XXL
- VAE
- Wan2.1 VAE compatible
Wan2.2 A14B T2V + ComfyUI frequently asked questions
Common questions about deploying Wan2.2 A14B T2V + ComfyUI on HexGrid Cloud.
Is Wan2.2 A14B a 14B model?
Not in total parameter count. It contains two approximately 14B experts, totaling roughly 27B parameters, with about 14B active at each denoising step.
Why are there two ComfyUI diffusion files?
One represents the high-noise expert used earlier in denoising and the other the low-noise expert used later to refine detail.
What resolutions are officially supported?
Wan lists both 480p and 720p for T2V-A14B.
Can it generate audio?
No. Wan2.2 T2V-A14B is a video-generation model; it does not jointly generate a native soundtrack like MiniMax H3.
Deploy Wan2.2 A14B T2V + ComfyUI today
$0.45 per hour on 1× RTX 4090, billed per minute, never shared.