Open Models on Your Own GPUs
Deploy Llama, Qwen, Mistral or DeepSeek on a dedicated GPU, billed by the hour instead of the token. No shared endpoints and no rate limits — talk to us about anything not listed.21 models · from $0.20/hr
Qwen3 Embedding 4B ↗
Alibaba
Qwen3 Reranker 4B ↗
Alibaba
BGE Reranker v2 Gemma ↗
BAAI
mxbai Rerank Large v2 ↗
Mixedbread
Qwen3.5 9B ↗
Alibaba
Llama 3.1 8B Instruct ↗
Meta
Ministral-3 8B Instruct ↗
Mistral AI
Qwen3 Embedding 8B ↗
Alibaba
Qwen3 Reranker 8B ↗
Alibaba
Wan 2.2 TI2V 5B ↗
Alibaba
Ministral-3 14B Instruct ↗
Mistral AI
MiniMax H3 Ref2V-A ↗
MiniMax
MiniMax H3 TI2V-A ↗
MiniMax
Devstral Small-2 24B Instruct ↗
Mistral AI
Gemma 4 31B IT ↗
Nemotron-3 Nano 30B A3B ↗
NVIDIA
Qwen3.6 27B ↗
Alibaba
Qwen3.5 27B ↗
Alibaba
Wan 2.2 T2V A14B ↗
Alibaba
Wan 2.2 I2V A14B ↗
Alibaba
Llama 3.3 70B Instruct ↗
Meta
Running something else?
Any open-weight model from Hugging Face, your own fine-tune, or a private checkpoint — point us at it and we will size the GPU.
Talk to usvRAM estimates cover weights at the selected precision plus ~20% for KV cache and activations. Video workflows use measured figures instead, so the precision toggle does not move them.
How model hosting works
We do not run a shared inference API. You pick a model, we put it on a GPU that is yours alone, and you talk to it over an OpenAI-compatible endpoint.
Single-tenant by default
The model runs on a GPU that belongs to you for as long as you keep it. No shared endpoint, no queue behind other tenants.
Priced by the hour, not the token
You pay for the GPU. Throughput, batch size and how hard you push it are yours to decide.
Any open weights, any GPU
Start from the catalogue or point us at a Hugging Face repo, then pick the card that fits your budget.
Your weights, your data
Prompts and completions stay inside your instance. Nothing is logged, sampled or used for training.
Pick the model, pick the card, keep the weights and the traffic to yourself.