Qwen3.8 27B Hardware Requirements: VRAM by Quant, and Why Context Costs So Little

A 24 GB card runs Qwen3.8 27B at Q4 or Q5. Q8 needs 32 GB. The surprise is the context cache: 128K tokens costs about 8 GB, a quarter of what an…

Aliteq
Voltage · Hardware Editor

The short answer

Qwen3.8 27B needs about 18.5 GB of VRAM at Q4_K_M and 30.3 GB at Q8_0 with 32K tokens of context, by our formula. That puts Q4 and Q5 on a 24 GB card, Q6 on the edge of one, and Q8 on a 32 GB card.…

Weights decide the card: Q4_K_M 15.7 GB, Q5_K_M 18.4, Q6_K 21.2, Q8_0 27.5, BF16 51.7

24 GB: Q4_K_M with about 80K tokens of context, or Q5_K_M with about 40K

32 GB (RTX 5090 class): Q6_K with about 108K tokens, or Q4_K_M with about 197K

Derived, not measured: our engine ran 2 to 4 percent above real GGUF file sizes

Aliteq

Read the full story

Qwen3.8 27B Hardware Requirements: VRAM by Quant, and Why Context Costs So Little

Read the full story on Aliteq