aliteq.

What it actually costs to run a model locally

Not “best GPU” lists. For each model: the memory it really needs at each quantisation, which cards fit it, and what the same job costs to rent by the hour — computed from the model's own config and live provider pricing.

40 models56 GPUs167 live prices tracked hourly

Interactive

Watch context length eat your GPU

Drag a slider and see the KV cache — not the weights — push a model off the card.

Live · hourly

Cloud GPU prices right now

Cheapest H100 and 5090 rental this hour, and the cheapest GPU that runs each model.

What fits at each VRAM tier

How many of the 40 tracked models run at a usable quant on each card class, at 16k context.

8GB9/40

RTX 4060 · 3050

12GB14/40

RTX 3060 · 4070

16GB23/40

RTX 5060 Ti · 4060 Ti

24GB27/40

RTX 3090 · 4090

48GB33/40

2× 3090 · RTX 6000

On a 24 GB card at 16k context

ModelParamsBest fitUses
Qwen3 235B-A22B (MoE)MoE235.1Bdoesn't fit—
Qwen3.5 122B-A10B (MoE)MoE125.1Bdoesn't fit—
gpt-oss 120B (MoE)MoE116.8Bdoesn't fit—
Llama 4 Scout 17B-16E (MoE)MoE108.6Bdoesn't fit—
GLM-4.5-Air (MoE)MoE106.0Bdoesn't fit—
Qwen3-Next 80B-A3B (MoE)MoE80.0Bdoesn't fit—
Qwen3-Coder-Next 80B-A3B (MoE)MoE79.7Bdoesn't fit—
DeepSeek R1 Distill 70B70.6Bdoesn't fit—
Llama 3.3 70B Instruct70.6Bdoesn't fit—
Mixtral 8x7B (MoE)MoE46.7BQ2_K
21 GB
Qwen3.5 35B-A3B (MoE)MoE36.0BQ4_K_M
21 GB
Qwen3.6 35B-A3B (MoE)MoE36.0BQ4_K_M
21 GB
Qwen2.5 Coder 32B32.8BQ4_K_M
23 GB
DeepSeek R1 Distill 32B32.8BQ4_K_M
23 GB
Qwen3 32B32.8BQ4_K_M
23 GB
Gemma 4 31B32.7BQ5_K_M
24 GB
Qwen3-Coder 30B-A3B (MoE)MoE30.5BQ5_K_M
22 GB
Qwen3 30B-A3B (MoE)MoE30.5BQ5_K_M
22 GB
Qwen3.8 27B27.8BQ6_K
23 GB
Qwen3.6 27B27.8BQ6_K
23 GB
Qwen3.5 27B27.8BQ6_K
23 GB
Gemma 3 27B27.4Bdoesn't fit—
Gemma 2 27B27.2Bdoesn't fit—
Gemma 4 26B-A4B (MoE)MoE26.5BQ6_K
21 GB
Mistral Small 24B23.6BQ6_K
21 GB
gpt-oss 20B (MoE)MoE20.9BQ8_0
22 GB
DeepSeek R1 Distill 14B14.8BQ8_0
18 GB
Qwen3 14B14.8BQ8_0
18 GB
Phi-4 14B14.7BQ8_0
18 GB
Gemma 3 12B12.2BQ8_0
14 GB
Gemma 4 12B12.0BF16
23 GB
Qwen3.5 9B9.7BF16
19 GB
Gemma 2 9B9.2Bdoesn't fit—
Qwen3 8B8.2BF16
18 GB
Llama 3.1 8B Instruct8.0Bdoesn't fit—
Gemma 4 E4B8.0BF16
16 GB
Qwen2.5 7B Instruct7.6BF16
16 GB
Mistral 7B Instruct v0.37.2BF16
16 GB
Gemma 4 E2B5.1BF16
10 GB
Llama 3.2 3B Instruct3.2BF16
9 GB

Hardware guides

The longer decisions — which card, rent or buy, what the marketing leaves out.

GLM-4.5-Air on local hardware

The consumer-runnable GLM: VRAM by quant, which cards and unified-memory boxes clear it, and what to expect from published numbers — the written companion to this page.

Cheapest way to run GLM in the cloud

Air on one 80 GB card, GLM-4.6 on two, GLM-5 on a pair of H200s — the multi-GPU maths with live, dated rental prices.

Run Mistral Small 24B locally

VRAM by quant (Q4 ~14 GB, Q8 ~25 GB, BF16 ~55 GB), which cards clear it, what tokens/s to expect from published benchmarks — and what a 24 GB card rents for when yours doesn't.

Cheapest RTX 4090 rental

Live per-hour 4090 prices on Vast and Runpod, the spot-vs-on-demand lever, and what a 24 GB card actually runs — the rent-a-card answer for everything in the 14B–32B class.

Cheapest A100 80GB rental

Live A100 80GB prices cheapest-first, when the A100 beats an H100 on price-per-job, and the 70B-class workloads it's the sweet spot for.

Cheapest H100 rental

Live H100 prices across Vast and Runpod with the capture date, SXM vs PCIe, and the per-job maths for 70B–120B models.

H200 rental cost

141 GB is capacity, not speed — when the H200 premium over an H100 is worth paying, with live and published rates side by side.

RTX 5090 vs two RTX 3090s

Two 3090s give you 48 GB but not double the speed for chat — llama.cpp's default multi-GPU mode makes the cards take turns. What the benchmarks actually show.

Cheapest way to serve Llama 70B

"Cheapest" flips on duty cycle and concurrency — and in Europe the electricity bill alone can approach the cost of just renting. Plus the config defaults that silently break a 16k deployment.

What --n-cpu-moe actually does

The trick that makes MoE models 5× faster — except it usually makes them slower, and the famous speedup only happens when the model didn't fit in the first place.

Is the DGX Spark worth it?

NVIDIA spent three months optimising it and token generation got slower — because 128 GB of memory on a 273 GB/s bus holds huge models and reads them slowly. Who should actually buy one.

Memory figures are computed from each model's own config.json, using bits-per-weight constants checked against real published quantised file sizes (median error 0.7%). Rental prices are pulled hourly from provider APIs. Full methodology is on each model's page.