Got an RTX 5090? The Best Local LLMs for 32GB of VRAM in 2026
32GB does not unlock 70B models. It buys better quants and a much longer memory for the 27B to 35B class. Five picks that fit a 5090, with the math…
Aliteq
Chiptune · Retro & Preservation Editor
The short answer
On a 32GB card like the RTX 5090, the best local LLM depends on the job, and every good pick is a 20B to 35B model. 32GB does not unlock 70B: Llama 3.3 70B needs about 43 GB at Q4. What the extra…
Coding agents: Qwen3.6 35B-A3B at Q5_K_M with 128K tokens, about 27.1 GB
Dense all-rounder: Qwen3.8 27B at Q6_K, room for about 108K tokens
Images and documents: Gemma 4 31B at Q5_K_M, or Gemma 4 26B-A4B at near-lossless Q8_0
Speed and full context: gpt-oss-20b in its native format, 128K tokens in about 19.6 GB
Aliteq
Read the full story
Got an RTX 5090? The Best Local LLMs for 32GB of VRAM in 2026