DeepSeek's 'cheap' AI model costs 3 cents a query in the cloud. running it yourself needs $10k of GPUs

the model that made headlines for being nearly free to rent turns out to be one of the most expensive things you could try to self-host

Aliteq
Lena Fischer · AI & Local Compute Editor

The short version

DeepSeek V4 Flash is a 284-billion-parameter mixture-of-experts model that only activates about 13B parameters per token — but all 284B still have to sit in memory somewhere.

The short version

The smallest realistic local build (IQ1_S-XL) is about 57GB. The developer-recommended quant, Q4_K_M-XL, is 163GB.

The short version

No single consumer GPU comes close — this is strictly a multi-GPU rig or a large-unified-memory Mac.

The short version

Context length adds on top of that: roughly 0.5GB of extra VRAM at 8K tokens, 2.2GB at 32K, and 8.8GB at 128K.

The short version

For almost everyone, DeepSeek's own cloud API at 3 cents a query is still cheaper than the hardware needed to self-host this specific model.

Should you self-host DeepSeek V4 Flash?

Only with a real data-residency requirement or genuinely heavy daily volume. For everyone else, the cloud API is the cheaper and faster path — this model's economics were built around DeepSeek's own…

Aliteq

Read the full story

DeepSeek's 'cheap' AI model costs 3 cents a query in the cloud. running it yourself needs $10k of GPUs

Read the full story on Aliteq