ALITEQ.

What it actually costs to run a model locally

Not “best GPU” lists. For each model: the memory it really needs at each quantisation, which cards fit it, and what the same job costs to rent by the hour — computed from the model's own config and live provider pricing.

18 models52 GPUs99 live prices tracked hourly

On a 24 GB card at 16k context

ModelParamsBest fitUses
gpt-oss 120B (MoE)MoE120.4Bdoesn't fit
Llama 3.3 70B Instruct70.6Bdoesn't fit
Mixtral 8x7B (MoE)MoE46.7BQ2_K21 GB
Qwen2.5 Coder 32B32.8BQ4_K_M23 GB
DeepSeek R1 Distill 32B32.8BQ4_K_M23 GB
Qwen3 32B32.8BQ4_K_M22 GB
Qwen3 30B-A3B (MoE)MoE30.5BQ5_K_M22 GB
Gemma 3 27B27.4Bdoesn't fit
Gemma 2 27B27.2Bdoesn't fit
Mistral Small 24B23.6BQ6_K22 GB
gpt-oss 20B (MoE)MoE21.5BQ8_023 GB
Qwen3 14B14.8BQ8_018 GB
Phi-4 14B14.7BQ8_018 GB
Gemma 2 9B9.2Bdoesn't fit
Qwen3 8B8.2BF1618 GB
Llama 3.1 8B Instruct8.0Bdoesn't fit
Qwen2.5 7B Instruct7.6BF1616 GB
Mistral 7B Instruct v0.37.2BF1616 GB

Memory figures are computed from each model's own config.json, using bits-per-weight constants checked against real published quantised file sizes (median error 0.7%). Rental prices are pulled hourly from provider APIs. Full methodology is on each model's page.