Private deployments

Open Models on Your Own GPUs

Deploy Llama, Qwen, Mistral or DeepSeek on a dedicated GPU, billed by the hour instead of the token. No shared endpoints and no rate limits — talk to us about anything not listed.

21 models · from $0.20/hr

Precision
Sort

vRAM estimates cover weights at the selected precision plus ~20% for KV cache and activations. Video workflows use measured figures instead, so the precision toggle does not move them.

How model hosting works

We do not run a shared inference API. You pick a model, we put it on a GPU that is yours alone, and you talk to it over an OpenAI-compatible endpoint.

Single-tenant by default

The model runs on a GPU that belongs to you for as long as you keep it. No shared endpoint, no queue behind other tenants.

Priced by the hour, not the token

You pay for the GPU. Throughput, batch size and how hard you push it are yours to decide.

Any open weights, any GPU

Start from the catalogue or point us at a Hugging Face repo, then pick the card that fits your budget.

Your weights, your data

Prompts and completions stay inside your instance. Nothing is logged, sampled or used for training.

Pick the model, pick the card, keep the weights and the traffic to yourself.