Kimi K3 just topped the coding leaderboard. here's the model you can actually run on one GPU that gets you most of the way

The best open coding model in the world needs a datacenter to run. The best coding model you can run on a single 24GB card is a different name…

Aliteq
Lena Fischer · AI & Local Compute Editor

The short answer

The best local coding model for a single 24GB GPU in 2026 is Qwen3-Coder-30B-A3B, a Mixture-of-Experts model that fits in ~18–20GB and runs fast because only ~3B of its 30B parameters are active per…

Qwen3-Coder-30B-A3B — best all-round local coding model: MoE, ~20GB at Q4, fast (~3B active params), strong multi-file coherence.

Qwen2.5-Coder-32B — dense alternative, ~20GB at Q4, thorough but slower per token than the MoE model.

24GB is the sweet spot — an RTX 3090 (cheapest), 4090, or 5090 all run these comfortably. 16GB cards are tight; 12GB fits only smaller 14B coders.

The MoE trick: 30B total parameters but only ~3B active per token means it computes like a small model while reasoning like a large one — fast *and* capable.

Kimi K3 is the ceiling you can't run — but a 30B coder gets most people most of the way for real coding tasks.

Aliteq

Read the full story

Kimi K3 just topped the coding leaderboard. here's the model you can actually run on one GPU that gets you most of the way

Read the full story on Aliteq