Got an RTX 5090? The Best Local LLMs for 32GB of VRAM in 2026

32GB does not unlock 70B models. It buys better quants and a much longer memory for the 27B to 35B class. Five picks that fit a 5090, with the math…

Aliteq
Chiptune · Retro & Preservation Editor

The short answer

On a 32GB card like the RTX 5090, the best local LLM depends on the job, and every good pick is a 20B to 35B model. 32GB does not unlock 70B: Llama 3.3 70B needs about 43 GB at Q4. What the extra…

Coding agents: Qwen3.6 35B-A3B at Q5_K_M with 128K tokens, about 27.1 GB

Dense all-rounder: Qwen3.8 27B at Q6_K, room for about 108K tokens

Images and documents: Gemma 4 31B at Q5_K_M, or Gemma 4 26B-A4B at near-lossless Q8_0

Speed and full context: gpt-oss-20b in its native format, 128K tokens in about 19.6 GB

Aliteq

Read the full story

Got an RTX 5090? The Best Local LLMs for 32GB of VRAM in 2026

Read the full story on Aliteq