only got 8GB of VRAM? you can still run real local AI — here's exactly what fits

8GB is the entry floor for local AI, and plenty of capable models run on it if you pick the right size and quantization. Here's what actually works,…

Aliteq
Lena Fischer · AI & Local Compute Editor

The 8GB reality

7–8B models at Q4 are your home — Llama 8B, Mistral 7B, Qwen3-8B run comfortably.

The 8GB reality

Keep context modest — long context eats VRAM; 4–8k is a safe range on 8GB.

The 8GB reality

Q4 is the sweet spot — smaller quants fit, but Q4 balances size and quality well.

The 8GB reality

Skip 13B+ dense models — they won't fit 8GB without heavy compromise.

The 8GB reality

MoE with offload can stretch further, but 8GB is genuinely a 7–8B tier card.

Aliteq

Read the full story

only got 8GB of VRAM? you can still run real local AI — here's exactly what fits

Read the full story on Aliteq