GLM-4.5-Air is 106B total but only 12B active, ~60GB at 4-bit — small enough for one 80GB card or a 128GB unified box. Exactly what hardware runs it, and the cheapest way if you don't own one.
This post contains affiliate links. If you buy through them, Aliteq may earn a commission — at no extra cost to you. Prices verified at publish time.
Share
GLM-4.5-Air is the GLM most people can actually run at home: 106 billion total parameters but only 12 billion active per token (it's a mixture-of-experts model), with a 128K context, per its Hugging Face model card. At 4-bit that's about 60GB of weights — small enough for a single big card or a 128GB unified-memory box. Here's exactly what hardware runs it, and the cheapest way if you don't own one. Sourced, no hands-on claims.
VRAM by quantisation
The numbers below are the weight footprint at each quant (add KV cache for long context; the 128K window costs extra memory on top). GLM-4.5-Air's 12B active parameters keep it responsive even when all 106B are resident.
GLM-4.5-Air — memory by quant · verified 24 Sep 2026
4-bit (Q4/INT4)
Approx VRAM
~60GB
Fits
80GB card / 96GB card / 128GB unified
Note
The practical local pick
FP8
Approx VRAM
~106GB
Fits
128GB unified / 2×80GB
Note
Higher fidelity
BF16 (full)
Approx VRAM
~212GB
Fits
Multi-GPU / cloud
Note
Rarely needed locally
Approx VRAM
Fits
Note
4-bit (Q4/INT4)
~60GB
80GB card / 96GB card / 128GB unified
The practical local pick
FP8
~106GB
128GB unified / 2×80GB
Higher fidelity
BF16 (full)
~212GB
Multi-GPU / cloud
Rarely needed locally
Which hardware to run it on
Three routes fit the ~60GB 4-bit build. A single 80GB data-centre card (A100 or H100) is the cleanest. A 96GB RTX PRO 6000 or 128GB unified-memory box — a Strix Halo mini-PC or a Mac Studio — holds it with room for a long context and is quieter and cheaper to idle. A pair of 48GB cards works too with tensor-parallelism. Match a specific setup with our best GPU for GLM-4.5-Air and cost-to-run pages.
No 80GB card? Rent one to run GLM-4.5-AirReferral link
About 60GB at 4-bit for the weights, per Unsloth's quants and the model card — so it fits a single 80GB card, a 96GB RTX PRO 6000, or a 128GB unified-memory machine, with the 128K context adding memory on top. FP8 is roughly 106GB. It's the most single-box-friendly GLM.
Can a Strix Halo or Mac Studio run GLM-4.5-Air?
Yes — a 128GB unified-memory machine holds the 4-bit build (~60GB) comfortably with room for context. Because only 12B parameters are active per token, throughput on unified memory is reasonable for a 106B model. It's one of the better uses of a big unified-memory mini-PC or Mac.
What's the cheapest way to run GLM-4.5-Air if I don't own the hardware?
Rent a single 80GB card in the cloud — an A100 80GB is about $0.47/hour on Vast.ai spot as of 24 September 2026. You pay only while it runs, so trying the model costs a dollar or two. See our cheapest way to run GLM in the cloud for the full breakdown.
GLM-4.5-Air is the sweet spot of the GLM family for local use: top-tier quality that fits one big card or a unified-memory box. See the best GPU for GLM-4.5-Air and cost to run, the GLM local pillar, and GLM vs Qwen3 if you're weighing a lighter model.