aliteq.

What it costs to run Llama 4 Scout 17B-16E (MoE)

Trying to decide what to buy? The best GPU for Llama 4 Scout 17B-16E (MoE) →

108.6B paramsMoE · 16 experts, 1 active10240k max contextllama4

Cheapest way to rent it right now

$0.11/hr

NVIDIA L40S · Q2_K · spot · ≈ $26/mo at 8h/day

Live price captured Sun, 04 Oct 2026 19:21:04 GMT. Referral link — we may earn a commission at no cost to you; it never changes which card is cheapest.

Memory needed, by quantisation

At 16k context. Weights + KV cache + ~0.8 GB overhead.

QuantWeightsKV cacheTotalQuality
F16202.4 GB0.8 GB203.9 GBFull precision. Reference quality, twice the size of Q8 for no practical gain in most chat use.
Q8_0107.5 GB0.8 GB109.0 GBEffectively lossless. Use when VRAM is not the constraint.
Q6_K83.0 GB0.8 GB84.5 GBVery close to Q8 at meaningfully less memory. A safe high-quality pick.
Q5_K_M71.8 GB0.8 GB73.4 GBSmall, hard-to-notice quality loss. Good balance.
Q5_070.1 GB0.8 GB71.6 GBOlder-style 5-bit. Q5_K_M is usually the better pick at the same size.
Q4_K_M61.3 GB0.8 GB62.9 GBThe community default. Best quality-per-gigabyte for most people.
Q4_057.5 GB0.8 GB59.1 GBOlder-style 4-bit, measurably worse than Q4_K_M at a similar size. Avoid unless required.
IQ4_XS53.8 GB0.8 GB55.3 GBNewer 4-bit, smaller than Q4_K_M with comparable quality. Needs a recent llama.cpp.
Q3_K_M49.5 GB0.8 GB51.0 GBNoticeable degradation. Use to fit a larger model that would otherwise not run.
Q2_K42.4 GB0.8 GB43.9 GBHeavy degradation. Almost always better to run a smaller model at Q4_K_M instead.
IQ3_XS41.7 GB0.8 GB43.3 GBAggressive. Usually better than a smaller model at Q4, but test before trusting it.

Cards that can run it

Best quantisation each card fits at 16k context, with live rental price where we track one.

GPUVRAMBest fitUses~tok/sRent from
Apple M3 Ultra512 GBF16204 GB——
Apple M2 Ultra192 GBQ8_0109 GB——
Apple M4 Max (16-core CPU / 40-core GPU)128 GBQ8_0109 GB——
AMD Ryzen AI Max+ 395 (Strix Halo, 128 GB)128 GBQ8_0109 GB——
NVIDIA DGX Spark (GB10 Grace Blackwell)128 GBQ8_0109 GB——
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition96 GBQ6_K84 GB—$1.690/hrRunpod
NVIDIA RTX PRO 6000 Blackwell Workstation Edition96 GBQ6_K84 GB——
Apple M2 Max96 GBQ6_K84 GB——
Apple M4 Pro64 GBQ4_K_Mtight63 GB——
NVIDIA RTX 6000 Ada Generation48 GBQ2_Ktight44 GB—$0.521/hrVast.ai
AMD Radeon PRO W790048 GBQ2_Ktight44 GB——
NVIDIA RTX A600048 GBQ2_Ktight44 GB—$0.330/hrRunpod

How these numbers are produced

  • Model shape comes from the model's own config.json — 12 layers, 40 attention heads, 8 KV heads.
  • Weight size is parameters × effective bits-per-weight. Those constants are checked against real published quantised file sizes — median error 0.7% across 40 measurements.
  • KV cache is 2 × layers × KV-heads × head-dim × context × 2 bytes. Using KV-heads rather than attention heads is what makes this correct for grouped-query attention; treating a GQA model as multi-head overstates the cache by up to 8×.
  • Tokens/sec is an estimate, not a benchmark. Generation is memory-bandwidth-bound, so this is bandwidth ÷ weight-bytes derated to 75%. Real throughput depends on your runtime, batch size and kernels. We don't run our own hardware tests — see our editorial policy.
  • This is a mixture-of-experts model. Generation reads only the experts active for each token (1 of 16), so throughput is scored on ~17.0B active parameters, not the full 109B. Memory is the opposite: every expert must still be resident in VRAM, so the figures above use the full weight set.Hybrid MoE: 48 layers, 3 of 4 use chunked attention (8,192-token chunks: 36 layers × 8 KV heads × 128, fixed ~1.2 GB at fp16) and every 4th is a global NoPE layer; `layers`/kv_heads/head_dim = the 12 global layers. Every layer is MoE: 16 routed experts, 1 routed + 1 shared per token; active = vendor-published 17B, but all ~109B weights (vision included) stay resident. Max context 10M per config. Official repo is gated: config.json read 4 Oct 2026 from the unsloth/Llama-4-Scout-17B-16E-Instruct mirror (safetensors total identical to the official API figure).
  • Some memory-bandwidth figures are not yet independently verified.Where that's the case we show no tokens/sec at all rather than a number we can't stand behind.
  • Rental prices are pulled hourly from provider APIs (last updated Sun, 04 Oct 2026 19:21:04 GMT). Spot/interruptible pricing can change or vanish without notice.

Try it interactively

Drag the context slider and watch the KV cache fill the card

Worth reading before you buy

Other models