Qwen2.5-Coder-32B locally: a private Copilot on a 24GB card

Qwen2.5-Coder-32B still runs a serious private coding assistant on a single 24GB card, with a 32K native context (128K with YaRN). The VRAM math,…

Aliteq
Lena Fischer · AI & Local Compute Editor

The short answer

Qwen2.5-Coder-32B is a strong open coding model that runs a private, offline Copilot on a single 24GB card. At Q4_K_M our VRAM engine computes about 23 GB — a tight but real fit for an RTX 3090 or…

32.5B params, dense, 32K native context (128K with YaRN), Apache-2.0 — from the model card (Nov 2024).

Q4_K_M ~23.3 GB (computed by our VRAM engine, 16K context) — a tight fit on a 24GB card.

Run qwen2.5-coder:32b explicitly; the newer Qwen3-Coder is the MoE option.

Aliteq

Read the full story

Qwen2.5-Coder-32B locally: a private Copilot on a 24GB card

Read the full story on Aliteq