Can't hold GLM locally? Rent it. GLM-4.5-Air needs one 80GB card (~$0.47/hr); GLM-5 needs two H200s (~$5.26/hr). The card maths and live dated prices for each version.
This post contains affiliate links. If you buy through them, Aliteq may earn a commission — at no extra cost to you. Prices verified at publish time.
Share
If your machine can't hold GLM, renting is the cheapest way to run it — and the maths is just 'how much memory does this quant need, and how many cards is that.' As of 24 Sep 2026 an 80GB card is about $0.47/hour (A100, Vast.ai spot) and an H200 about $2.63/hour, so even GLM-5 on a multi-GPU rental is a few dollars an hour. Here's exactly what to rent for each GLM, with the card maths shown. Part of our cloud-GPU pricing pillar.
What to rent for each GLM
The rule is simple: take the quant's memory footprint, divide by the card's memory, round up. Prices below are the cheapest live offers from our tracker (Vast.ai spot, Runpod on-demand), captured 24 Sep 2026; multi-GPU rates are per-card × count.
Cheapest cloud setup per GLM · captured 24 Sep 2026
GLM-4.5-Air (~60GB)
Cards needed
1× A100 80GB
Vast.ai spot
$0.47/hr
Runpod on-demand
$1.19/hr
GLM-4.6 (~135GB)
Cards needed
2× H100 80GB or 1× H200
Vast.ai spot
~$3.58/hr (2×H100)
Runpod on-demand
$3.59/hr (1×H200)
GLM-5 (~241GB)
Cards needed
2× H200 (~282GB)
Vast.ai spot
~$5.26/hr
Runpod on-demand
~$7.18/hr
Cards needed
Vast.ai spot
Runpod on-demand
GLM-4.5-Air (~60GB)
1× A100 80GB
$0.47/hr
$1.19/hr
GLM-4.6 (~135GB)
2× H100 80GB or 1× H200
~$3.58/hr (2×H100)
$3.59/hr (1×H200)
GLM-5 (~241GB)
2× H200 (~282GB)
~$5.26/hr
~$7.18/hr
Rent an 80GB card for GLM-4.5-Air from ~$0.47/hrReferral link
Vast.ai's spot market is cheapest but interruptible — fine for GLM-4.5-Air inference you can checkpoint. For GLM-4.6 and GLM-5 you need two to four cards at once, and multi-GPU spot capacity from a single host is rarer, so an on-demand multi-GPU instance (or a high-reliability host) is usually more practical — budget accordingly. Either way you only pay while the box is on, so a few hours with GLM-5 costs a handful of dollars, not a hardware purchase.
Referral link
Rent the cheapest card that fits your GLM
An 80GB A100 from ~$0.47/hr for GLM-4.5-Air, up to a 2×H200 rig from ~$5.26/hr for GLM-5 (Vast.ai spot, 24 Sep 2026). Checkpoint on spot; you pay only while it runs.
Referral link — we may earn a commission at no cost to you. Prices on our compare page are the provider's live figures, cheapest first; this never changes the ranking.
Frequently asked
What's the cheapest way to run GLM in the cloud?
Match the quant to the fewest cards. GLM-4.5-Air (~60GB at 4-bit) needs one 80GB card — about $0.47/hour on Vast.ai spot (24 Sep 2026). GLM-4.6 (~135GB) needs two 80GB cards or one H200; GLM-5 (~241GB) needs two H200s or four 80GB cards. Spot is cheapest but interruptible, so checkpoint; you pay only while the instance runs.
How many GPUs do I need to run GLM-5 in the cloud?
GLM-5's 2-bit quant is about 241GB, so you need roughly 282GB of GPU memory — two H200s (141GB each) or four 80GB cards (A100/H100). Multi-GPU spot capacity is limited, so an on-demand multi-GPU instance is often the practical choice. At Vast.ai spot rates that's roughly $5.26/hour for a 2×H200 setup.
Is it cheaper to rent or buy hardware for GLM?
For occasional use, rent — an evening with GLM costs a few dollars versus a large-memory machine costing thousands. Buying a 128GB unified-memory box makes sense only if you run GLM-4.5-Air constantly; for the bigger versions, renting multi-GPU by the hour avoids owning hardware you'd rarely fill. See our rent-vs-buy math.
Referral link
Managed pods + serverless
Need a stable multi-GPU box for GLM-4.6 or GLM-5? Runpod on-demand avoids spot reclaims — H200 about $3.59/hr, A100 80GB $1.19/hr (24 Sep 2026).