Qwen3-30B-A3B fits a 24GB card at ~17.5GB; GLM-4.5-Air needs ~60GB. GLM ranks a touch higher; Qwen3 runs on far less. How to choose, with sourced numbers.
This post contains affiliate links. If you buy through them, Aliteq may earn a commission — at no extra cost to you. Prices verified at publish time.
Share
GLM and Qwen3 are the two Chinese open-weight families at the top of the 2026 rankings, and for local use they pull in different directions. Qwen3 has a genuinely tiny-footprint MoE — Qwen3-30B-A3B runs in about 17.5GB — while the most local GLM, GLM-4.5-Air, needs about 60GB at 4-bit. GLM tends to rank higher on quality; Qwen3 fits a single consumer card. Here's how to choose, with sourced numbers and no hands-on claims.
The local footprint, side by side
GLM vs Qwen3 — local memory · verified 24 Sep 2026
Qwen3-30B-A3B
Params (total / active)
30B / 3B
Smallest practical VRAM
~17.5GB
Single-card home
One 24GB card (3090/4090)
GLM-4.5-Air
Params (total / active)
106B / 12B
Smallest practical VRAM
~60GB (4-bit)
Single-card home
One 80GB card / 128GB unified
GLM-5 (flagship)
Params (total / active)
744B / 40B
Smallest practical VRAM
~241GB (2-bit)
Single-card home
256GB Mac / multi-GPU / cloud
Params (total / active)
Smallest practical VRAM
Single-card home
Qwen3-30B-A3B
30B / 3B
~17.5GB
One 24GB card (3090/4090)
GLM-4.5-Air
106B / 12B
~60GB (4-bit)
One 80GB card / 128GB unified
GLM-5 (flagship)
744B / 40B
~241GB (2-bit)
256GB Mac / multi-GPU / cloud
The practical read: if you have a 24GB card, Qwen3-30B-A3B just works, and it's excellent — see our best GPU for Qwen3 guide. If you have an 80GB card or a 128GB unified-memory box and want a notch more quality, GLM-4.5-Air is the pick. And if you want GLM-5's frontier quality, that's unified-memory-or-cloud territory. Match any of them to a card with our cost-to-run tool.
Want to try GLM-4.5-Air without an 80GB card?Referral link
It depends on your hardware. Qwen3-30B-A3B runs in ~17.5GB on a single 24GB card, so it's the easier local model; GLM-4.5-Air needs ~60GB but ranks a little higher on quality, and GLM's flagship ties Kimi K3 at the top of the open-weight index. Pick Qwen3 for modest hardware, GLM when you have more memory and want the extra quality.
Can I run GLM on a 24GB card like a Qwen3 model?
Not the same way. Qwen3-30B-A3B fits a 24GB card directly at ~17.5GB; the smallest GLM (4.5-Air) needs ~60GB at 4-bit, so on a single 24GB card you'd need heavy MoE offload to system RAM, which is slow. For a 24GB card, Qwen3-30B-A3B is the far better fit.
Which has the longer context?
Both are generous: Qwen3 supports very long contexts (256K natively, extendable further), and GLM-4.5-Air is 128K with GLM-5 at 200K. For most local use either is more than enough; longer contexts cost extra memory on top of the weights.