RunPod Serverless pricing, explained: why it looks pricier but usually isn't

Serverless per-hour rates look higher than a pod — but they bill per second and scale to zero. The whole decision is one number: how busy your GPU…

Aliteq
Ravi Malhotra · Hardware Editor

The short version

Billed per second, scales to zero — you pay only while a worker is actively processing, not for idle time.

The short version

Flex rate card (RunPod published): 24GB ~$0.69/hr, 48GB ~$1.22/hr, A100 ~$2.72/hr, H100 ~$4.79/hr, B200 ~$8.64/hr.

The short version

Per-active-hour it's pricier than a pod (H100 serverless ~$4.79 vs ~$1.99 pod, 22 Sep 2026) — the saving is only paying for active seconds.

The short version

Break-even ≈ 10 active hours/day for an H100: busier than that, a pod is cheaper; quieter, serverless wins.

The short version

Watch cold starts + storage — you're billed during worker init, and idle volumes cost ~$0.10–$0.20/GB/month.

Aliteq

Read the full story

RunPod Serverless pricing, explained: why it looks pricier but usually isn't

Read the full story on Aliteq