aliteq.

GLM vs Qwen3 for local use: which top open model to actually run

Qwen3-30B-A3B fits a 24GB card at ~17.5GB; GLM-4.5-Air needs ~60GB. GLM ranks a touch higher; Qwen3 runs on far less. How to choose, with sourced numbers.

TensorUpdated Sep 247 min readWeb story
Flat illustration of a person comparing GLM and Qwen3 models for local use, on a teal background

This post contains affiliate links. If you buy through them, Aliteq may earn a commission — at no extra cost to you. Prices verified at publish time.

Share

GLM and Qwen3 are the two Chinese open-weight families at the top of the 2026 rankings, and for local use they pull in different directions. Qwen3 has a genuinely tiny-footprint MoE — Qwen3-30B-A3B runs in about 17.5GB — while the most local GLM, GLM-4.5-Air, needs about 60GB at 4-bit. GLM tends to rank higher on quality; Qwen3 fits a single consumer card. Here's how to choose, with sourced numbers and no hands-on claims.

The local footprint, side by side

GLM vs Qwen3 — local memory · verified 24 Sep 2026

Qwen3-30B-A3B

Params (total / active)
30B / 3B
Smallest practical VRAM
~17.5GB
Single-card home
One 24GB card (3090/4090)

GLM-4.5-Air

Params (total / active)
106B / 12B
Smallest practical VRAM
~60GB (4-bit)
Single-card home
One 80GB card / 128GB unified

GLM-5 (flagship)

Params (total / active)
744B / 40B
Smallest practical VRAM
~241GB (2-bit)
Single-card home
256GB Mac / multi-GPU / cloud

The practical read: if you have a 24GB card, Qwen3-30B-A3B just works, and it's excellent — see our best GPU for Qwen3 guide. If you have an 80GB card or a 128GB unified-memory box and want a notch more quality, GLM-4.5-Air is the pick. And if you want GLM-5's frontier quality, that's unified-memory-or-cloud territory. Match any of them to a card with our cost-to-run tool.

Vast.ai

Want to try GLM-4.5-Air without an 80GB card?Referral link

Rent a GPU for GLM

Frequently asked

Is GLM or Qwen3 better for local use?
It depends on your hardware. Qwen3-30B-A3B runs in ~17.5GB on a single 24GB card, so it's the easier local model; GLM-4.5-Air needs ~60GB but ranks a little higher on quality, and GLM's flagship ties Kimi K3 at the top of the open-weight index. Pick Qwen3 for modest hardware, GLM when you have more memory and want the extra quality.
Can I run GLM on a 24GB card like a Qwen3 model?
Not the same way. Qwen3-30B-A3B fits a 24GB card directly at ~17.5GB; the smallest GLM (4.5-Air) needs ~60GB at 4-bit, so on a single 24GB card you'd need heavy MoE offload to system RAM, which is slow. For a 24GB card, Qwen3-30B-A3B is the far better fit.
Which has the longer context?
Both are generous: Qwen3 supports very long contexts (256K natively, extendable further), and GLM-4.5-Air is 128K with GLM-5 at 200K. For most local use either is more than enough; longer contexts cost extra memory on top of the weights.

Qwen3 for a single consumer card, GLM for more memory and top-end quality — that's the split. See best GPU for Qwen3, the GLM local pillar, the GLM-4.5-Air hardware guide, and GLM vs DeepSeek V4.

Found this useful? Share it

Share
Tensor

Local AI & Automation Editor

Tensor

I'm US-based, I run more models at home than I'll admit to, and I've quantized more than I've finished reading about. I write about running AI on your own hardware and, lately, about what it costs a company to do the same — tokens per day, GPUs per month, and the GDPR questions nobody's sales deck answers.

Work out the hardware

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading