Serverless per-hour rates look higher than a pod — but they bill per second and scale to zero. The whole decision is one number: how busy your GPU actually is.
This post contains affiliate links. If you buy through them, Aliteq may earn a commission — at no extra cost to you. Prices verified at publish time.
Share
RunPod Serverless confuses people because its per-hour rates look higher than renting a plain pod — an H100 serverless worker is around $4.79/hour versus about $1.99/hour for an H100 pod. The trick is that serverless bills per second of active work and scales to zero when idle, so a lightly-used endpoint pays for minutes a day instead of 24 hours. Whether that's cheaper comes down to one number: how busy your GPU actually is. Here's how RunPod Serverless pricing works, the full rate card, and the utilisation break-even that decides it. This is a spoke of our cloud-GPU pricing pillar.
How serverless pricing actually works
A normal pod is a GPU you rent by the hour whether or not you're using it — the meter runs 24/7 until you shut it down. Serverless flips that: you deploy an endpoint, RunPod spins up 'workers' only when requests arrive, bills you per second from worker start to stop, and scales back to zero when the queue empties. Flex workers are pure pay-per-use (the rate card below); active workers stay warm for a discount and suit a steady baseline of traffic. The billing unit is roughly $0.00126/second on the mid tiers, rounded up — so an endpoint serving a hundred short requests a day costs cents, not a day's worth of GPU time.
RunPod Serverless (flex) — published rates, $/hr equivalent
16GB
Worker tier
$0.58
Rate ($/hr)
—
24GB (4090-class)
Worker tier
$0.69
Rate ($/hr)
$0.34 pod
48GB (A6000/L40)
Worker tier
$1.22
Rate ($/hr)
$0.69–0.79 pod
A100 80GB
Worker tier
$2.72
Rate ($/hr)
$1.19 pod
H100 80GB
Worker tier
$4.79
Rate ($/hr)
$1.99 pod
B200
Worker tier
$8.64
Rate ($/hr)
$5.98 pod
Worker tier
Rate ($/hr)
Pod equivalent (22 Sep 2026)
16GB
$0.58
—
24GB (4090-class)
$0.69
$0.34 pod
48GB (A6000/L40)
$1.22
$0.69–0.79 pod
A100 80GB
$2.72
$1.19 pod
H100 80GB
$4.79
$1.99 pod
B200
$8.64
$5.98 pod
The break-even: utilisation decides everything
Because serverless costs more per active hour but nothing when idle, the whole decision is how many hours a day your GPU is genuinely busy. Take the H100: a pod at ~$1.99/hour runs $47.76 a day whether you use it or not. A serverless H100 at ~$4.79/hour only bills active time, so it matches the pod's daily cost at about 10 active hours a day. Below that — a demo, an internal tool, bursty traffic, anything that sits idle most of the day — serverless is dramatically cheaper. Above it — a busy production endpoint serving requests most of the day — a plain pod (or a reserved/active worker) wins. Estimate your real active hours honestly; most side-project endpoints are nowhere near 10 hours of genuine GPU work a day.
Referral link
Deploy a scale-to-zero endpoint on RunPod Serverless
For bursty or low-volume inference, serverless bills only the seconds you use — an idle endpoint costs nothing. If your GPU would be busy more than ~10 hours a day, rent a pod instead. Check the live rate before you deploy.
Referral link — we may earn a commission at no cost to you. Prices on our compare page are the provider's live figures, cheapest first; this never changes the ranking.
The gotchas that inflate the bill
Cold starts are billed. You pay the per-second rate while a worker initialises (loading your container and model), before it does any work — big models with slow cold starts waste real money on every scale-up. Keep an active/warm worker if requests are frequent.
The idle-timeout trap. A flex worker doesn't stop the instant a request finishes — it stays alive for your configured idle timeout (waiting for the next request), and you're billed for every one of those idle seconds. Set it too high to dodge cold starts and a 'scale-to-zero' endpoint quietly bills like an always-on one; tune it to your real traffic gaps.
Idle storage isn't free. Container and volume disk run about $0.10/GB/month while running and $0.20/GB/month while stopped; network storage is ~$0.05–0.14/GB/month. A large model volume left attached adds up even at zero traffic.
Rounding up. Per-second billing rounds up, so very short, very frequent requests carry a little overhead each.
The pod is sometimes just simpler. For steady load, a pod at the tracker rate is cheaper and has no cold-start tax — don't reach for serverless out of habit.
Frequently asked
Is RunPod Serverless cheaper than a pod?
It depends entirely on utilisation. Serverless costs more per active hour (an H100 is around $4.79/hour serverless versus about $1.99/hour as a pod, 22 Sep 2026) but bills only while it's working and scales to zero when idle. For an H100 the break-even is roughly 10 active hours a day: busier than that, a pod is cheaper; quieter, serverless wins because you're not paying for idle time. Most low-volume or bursty endpoints are far cheaper on serverless.
How is RunPod Serverless billed?
Per second, from when a worker starts to when it stops, rounded up — including the cold-start initialisation time. Flex workers are pure pay-per-use; active workers stay warm for a discount and suit steady traffic. On top of compute you pay for storage: container and volume disk at roughly $0.10/GB/month running and $0.20/GB/month stopped, plus network storage. Always check RunPod's live pricing page for current figures.
What does a serverless H100 cost on RunPod?
Around $4.79/hour equivalent as a flex worker (RunPod's published rate — verify on their pricing page), billed per second of active work. Because it scales to zero, a lightly-used H100 endpoint can cost a few dollars a month, while the same card as an always-on pod would run about $1,400 a month. The serverless premium only bites once the endpoint is busy most of the day.
When should I use a pod instead of serverless?
Use a pod for sustained, predictable load — a training run, batch processing, or an endpoint that's busy more than about 10 hours a day — because the lower per-hour pod rate beats serverless once utilisation is high, and you avoid the cold-start tax. Use serverless for intermittent, bursty, or low-volume inference where the GPU would otherwise sit idle most of the time.
Serverless is a utilisation bet: it's cheaper when your GPU is idle most of the day and pricier when it's busy. Work out your real active hours, then choose. Compare live pod prices on our cloud-GPU compare page, read the pricing pillar, and see the cheapest H100 rental and cheapest cloud GPU overall for the pod side of the decision.