how to run GLM locally in 2026: which version, which GPU

GLM ties Kimi K3 at the top of the open-weight rankings, but 'run it locally' depends on which GLM — from the 106B Air that fits one card to the…

Aliteq
Lena Fischer · AI & Local Compute Editor

The short answer

GLM is a family of mixture-of-experts models from Zhipu. As of 24 September 2026: GLM-4.5-Air is 106B total / 12B active (128K context) and fits a single 80GB card or a 128GB unified-memory box at…

GLM-4.5-Air: 106B / 12B active, 128K context — ~60GB at 4-bit, the single-box local pick.

GLM-4.6: 355B — Unsloth 2-bit ~135GB; needs ~205GB RAM with MoE offload for usable speed.

GLM-5: 744B / 40B active, 200K context — 2-bit ~241GB (256GB Mac, big rig, or cloud).

MoE: all parameters load, but only the active set runs per token — fast for the size.

No big-memory box? Rent a card that fits, from cents an hour.

Aliteq

Read the full story

how to run GLM locally in 2026: which version, which GPU

Read the full story on Aliteq