aliteq.

the cheapest way to run GLM in the cloud (by version)

Can't hold GLM locally? Rent it. GLM-4.5-Air needs one 80GB card (~$0.47/hr); GLM-5 needs two H200s (~$5.26/hr). The card maths and live dated prices for each version.

Ravi MalhotraUpdated 1h ago7 min readWeb story
Flat illustration of a person renting cloud GPUs to run a large GLM model, on a teal background

This post contains affiliate links. If you buy through them, Aliteq may earn a commission — at no extra cost to you. Prices verified at publish time.

Share

If your machine can't hold GLM, renting is the cheapest way to run it — and the maths is just 'how much memory does this quant need, and how many cards is that.' As of 24 Sep 2026 an 80GB card is about $0.47/hour (A100, Vast.ai spot) and an H200 about $2.63/hour, so even GLM-5 on a multi-GPU rental is a few dollars an hour. Here's exactly what to rent for each GLM, with the card maths shown. Part of our cloud-GPU pricing pillar.

What to rent for each GLM

The rule is simple: take the quant's memory footprint, divide by the card's memory, round up. Prices below are the cheapest live offers from our tracker (Vast.ai spot, Runpod on-demand), captured 24 Sep 2026; multi-GPU rates are per-card × count.

Cheapest cloud setup per GLM · captured 24 Sep 2026

GLM-4.5-Air (~60GB)

Cards needed
1× A100 80GB
Vast.ai spot
$0.47/hr
Runpod on-demand
$1.19/hr

GLM-4.6 (~135GB)

Cards needed
2× H100 80GB or 1× H200
Vast.ai spot
~$3.58/hr (2×H100)
Runpod on-demand
$3.59/hr (1×H200)

GLM-5 (~241GB)

Cards needed
2× H200 (~282GB)
Vast.ai spot
~$5.26/hr
Runpod on-demand
~$7.18/hr
Vast.ai

Rent an 80GB card for GLM-4.5-Air from ~$0.47/hrReferral link

See live A100 prices

Spot vs on-demand, and the multi-GPU catch

Vast.ai's spot market is cheapest but interruptible — fine for GLM-4.5-Air inference you can checkpoint. For GLM-4.6 and GLM-5 you need two to four cards at once, and multi-GPU spot capacity from a single host is rarer, so an on-demand multi-GPU instance (or a high-reliability host) is usually more practical — budget accordingly. Either way you only pay while the box is on, so a few hours with GLM-5 costs a handful of dollars, not a hardware purchase.

Vast.aiReferral link

Rent the cheapest card that fits your GLM

An 80GB A100 from ~$0.47/hr for GLM-4.5-Air, up to a 2×H200 rig from ~$5.26/hr for GLM-5 (Vast.ai spot, 24 Sep 2026). Checkpoint on spot; you pay only while it runs.

Referral link — we may earn a commission at no cost to you. Prices on our compare page are the provider's live figures, cheapest first; this never changes the ranking.

Frequently asked

What's the cheapest way to run GLM in the cloud?
Match the quant to the fewest cards. GLM-4.5-Air (~60GB at 4-bit) needs one 80GB card — about $0.47/hour on Vast.ai spot (24 Sep 2026). GLM-4.6 (~135GB) needs two 80GB cards or one H200; GLM-5 (~241GB) needs two H200s or four 80GB cards. Spot is cheapest but interruptible, so checkpoint; you pay only while the instance runs.
How many GPUs do I need to run GLM-5 in the cloud?
GLM-5's 2-bit quant is about 241GB, so you need roughly 282GB of GPU memory — two H200s (141GB each) or four 80GB cards (A100/H100). Multi-GPU spot capacity is limited, so an on-demand multi-GPU instance is often the practical choice. At Vast.ai spot rates that's roughly $5.26/hour for a 2×H200 setup.
Is it cheaper to rent or buy hardware for GLM?
For occasional use, rent — an evening with GLM costs a few dollars versus a large-memory machine costing thousands. Buying a 128GB unified-memory box makes sense only if you run GLM-4.5-Air constantly; for the bigger versions, renting multi-GPU by the hour avoids owning hardware you'd rarely fill. See our rent-vs-buy math.
RunpodReferral link

Managed pods + serverless

Need a stable multi-GPU box for GLM-4.6 or GLM-5? Runpod on-demand avoids spot reclaims — H200 about $3.59/hr, A100 80GB $1.19/hr (24 Sep 2026).

Renting scales with the GLM you pick: one card for Air, a small rig for the big ones. Compare live prices on our cloud-GPU compare page, read the GLM local pillar and the pricing pillar, and see the cheapest way to run Llama 4 in the cloud for the same maths on a different model.

Found this useful? Share it

Share
Ravi Malhotra

Hardware Editor

Ravi Malhotra

Ravi has been building and taking apart PCs since the single-core days — his idea of a good weekend is a repaste and a spreadsheet full of thermals. He covers GPUs, CPUs and the build decisions that actually move frame rates, and he'd rather hand you a benchmark than a press release.

Work out the hardware

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading