aliteq.

the cheapest way to run Llama 4 in the cloud: Scout on one H100, Maverick only if you must

Scout fits a single H100 from ~$1.73/hr. Maverick needs 243GB — four H100s or two H200s. Here's the cheapest cloud setup for each, dated to real prices.

Lena FischerUpdated 1h ago7 min readWeb story

This post contains affiliate links. If you buy through them, Aliteq may earn a commission — at no extra cost to you. Prices verified at publish time.

Share

Renting is the sane way to try Llama 4 before you buy anything, because the two open models sit on opposite sides of a hardware cliff. Scout (109B total, 17B active) fits on a single H100 at Q4 — so cloud cost starts around $1.73/hour on Vast.ai spot or $1.99/hour on RunPod on-demand (captured 22 Sep 2026). Maverick (400B total) needs about 243GB of VRAM at Q4 — four H100s or two H200s — so it's roughly $5–8/hour and a different conversation. Here's the cheapest way to run each in the cloud, and when you shouldn't bother. This is a spoke of our cloud-GPU pricing pillar.

Scout: one H100 does it

Scout is the one most people actually want, and the good news is it fits a single 80GB card. At Q4_K_M it lands around 55–61GB, which leaves enough room on an H100 80GB for context. That makes the cheapest sensible cloud setup a single H100 — Vast.ai spot from ~$1.73/hour, or RunPod on-demand at ~$1.99/hour if you'd rather it not get reclaimed. If you want headroom for a long context window, an H200 (141GB) at ~$2.63/hour spot is the comfortable step up. All figures captured 22 Sep 2026.

Cheapest cloud setup for Llama 4 · $/hr · captured 22 Sep 2026

1× RTX 5090 32GB

Setup
$0.40
Cheapest (Vast spot)
$0.69
Stable (RunPod on-demand)
Scout @ ~1.78-bit (tight)

1× H100 80GB

Setup
$1.73
Cheapest (Vast spot)
$1.99
Stable (RunPod on-demand)
Scout @ Q4 — the sweet spot

1× H200 141GB

Setup
$2.63
Cheapest (Vast spot)
$3.59
Stable (RunPod on-demand)
Scout @ Q4 + long context

2× H200 (~282GB)

Setup
~$5.26
Cheapest (Vast spot)
~$7.18
Stable (RunPod on-demand)
Maverick @ Q4

4× H100 (~320GB)

Setup
~$6.92
Cheapest (Vast spot)
~$7.96
Stable (RunPod on-demand)
Maverick @ Q4
Vast.aiReferral link

Run Llama 4 Scout on a single H100 from ~$1.73/hr

Scout's ~55–61GB Q4 footprint fits one H100 80GB — the cheapest is Vast.ai spot as of 22 Sep 2026. Checkpoint if you're on spot; an evening of Scout costs a few dollars.

Referral link — we may earn a commission at no cost to you. Prices on our compare page are the provider's live figures, cheapest first; this never changes the ranking.

Maverick: the 243GB problem

Maverick is where the cost story changes. Its Q4 GGUF is about 243GB, so you need four H100s or two H200s just to load it. On Vast.ai spot that's roughly $6.92/hour for 4×H100 or about $5.26/hour for 2×H200 — the H200 pair is both cheaper and simpler. But multi-GPU nodes are rarer on the spot market and harder to keep alive interruptibly, so in practice you'll often want an on-demand multi-GPU instance or an enterprise neocloud, which pushes the real cost higher. For most people, Maverick in the cloud is a 'rent it for a specific job, then kill it' proposition, not something to leave running.

Runpod

Want a stable multi-GPU box for Maverick? RunPod on-demand.Referral link

See RunPod GPUs →

Frequently asked

What's the cheapest way to run Llama 4 Scout in the cloud?
A single H100 80GB. Scout quantised to Q4_K_M is about 55–61GB, which fits one H100 with room for context. As of 22 Sep 2026 that's roughly $1.73/hour on Vast.ai's spot marketplace or $1.99/hour on RunPod on-demand. If you accept an aggressive ~1.78-bit quant (about 24GB) you can squeeze Scout onto a single RTX 5090 at ~$0.40/hour spot, but expect a quality drop.
How much VRAM does Llama 4 Maverick need in the cloud?
About 243GB at Q4, because a mixture-of-experts model must load all 400B parameters into memory even though only 17B are active per token. That means four H100 80GB cards or two H200 141GB cards. On Vast.ai spot the 2×H200 route is the cheaper of the two at roughly $5.26/hour, though multi-GPU spot availability is limited, so on-demand is often more practical.
Is it cheaper to run Llama 4 in the cloud or buy hardware?
For occasional use, the cloud wins easily — an evening with Scout on a rented H100 costs a few dollars versus thousands for a card that fits it. Buying only pays off with heavy, sustained use, and Maverick-class hardware (multiple 80GB+ cards) is expensive to own. See our rent-vs-buy break-even math, and rent first to learn your real usage.
Can I run Llama 4 Scout on a cheaper GPU than an H100?
Only with aggressive quantisation. A ~1.78-bit dynamic GGUF shrinks Scout to around 24GB, which fits an RTX 5090 (32GB) or, tightly, an RTX 4090 (24GB) — both under $0.70/hour to rent. You trade output quality for the lower cost, so it's fine for experimentation but not the setup you'd serve from. At full Q4 quality, the H100 is the floor.

The cheap path is simple: Scout on a single H100 (from ~$1.73/hr spot), Maverick only when you specifically need it. Compare live prices on our cloud-GPU compare page, read the pricing pillar, and for the local side see how to run Llama 4 locally, the best GPU for Scout, Scout vs Maverick, and running Scout on unified memory.

Found this useful? Share it

Share
Lena Fischer

AI & Local Compute Editor

Lena Fischer

Lena runs more GPUs at home than she'll admit to and has quantized more models than she's finished reading about. She writes about running AI on your own hardware — what actually fits, what's genuinely fast, and what the polished cloud demos quietly leave out.

Work out the hardware

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading