the cheapest way to run Llama 4 in the cloud: Scout on one H100, Maverick only if you must

Scout fits a single H100 from ~$1.73/hr. Maverick needs 243GB — four H100s or two H200s. Here's the cheapest cloud setup for each, dated to real…

Aliteq
Lena Fischer · AI & Local Compute Editor

The short version

Scout (109B MoE): ~55–61GB at Q4_K_M — fits ONE H100 80GB. Cheapest: Vast.ai spot ~$1.73/hr, RunPod on-demand ~$1.99/hr (22 Sep 2026).

The short version

Maverick (400B MoE): ~243GB at Q4 — needs 4×H100 (~$7/hr spot) or 2×H200 (~$5.3/hr spot). Big jump.

The short version

MoE loads ALL parameters even though only 17B are active per token — that's why Scout needs 55GB+, not 17B-worth.

The short version

Cheap-and-dirty: an aggressive 1.78-bit dynamic quant shrinks Scout to ~24GB → a single RTX 5090 (~$0.40/hr spot), with a real quality trade-off.

The short version

Rent before you buy — Scout for an evening costs a few dollars; see the rent-vs-buy math.

For Maverick, price out the API first

If you just want Maverick's answers rather than control over the weights, a hosted API is often cheaper than renting 4×H100 by the hour — you pay per token instead of for idle GPU time. Rent your…

Aliteq

Read the full story

the cheapest way to run Llama 4 in the cloud: Scout on one H100, Maverick only if you must

Read the full story on Aliteq