Renting a GPU for an open model looks cheaper than paying per token until you price the same model on the cheapest API. We did it for four open models, from Llama 3.1 8B to Llama 3.3 70B, with live GPU rates from our own tracker and list prices read on 2 October 2026. Against the same model's API, one rented GPU almost never wins. Against a frontier model, it wins from about 2 to 6 million tokens a day.
This post contains affiliate links. If you buy through them, Aliteq may earn a commission — at no extra cost to you. Prices verified at publish time.
Share
"Should we run our own model?" usually gets answered with a GPU's hourly price and a frontier model's per-token price. That comparison flatters the GPU. The fair test is the same open model on someone else's GPU, billed per token, and those hosts are cheap.
So we priced both sides for four open models of different sizes. The GPU side uses aliteq's own tracker, which has logged Runpod's and Vast's prices every hour since 21 July 2026; this snapshot was captured at 19:21 UTC on 2 October 2026. For the API side we read each vendor's pricing page on 2 October 2026. The throughput figures are other people's published benchmarks. We didn't deploy or benchmark any of these models ourselves.
The break-even, model by model
One rented GPU beats the cheapest same-model API only if you push tens of millions of tokens through it every workday, and for three of our four models that's more than the card can serve. Against closed frontier models, break-even drops to between 2 and 11 million tokens a workday.
Millions of tokens a workday at which one rented GPU in business hours costs the same as the API. Above one GPU's ceiling, it never breaks even. · aliteq research
Break-even, millions of tokens per workday (one rented GPU, 8 hours a day)
Llama 3.1 8B on an RTX 4090 ($128 a month)
Cheapest same-model API
262M (never)
gpt-6-luna
40M
Claude Haiku 4.5
4.0M
gpt-6.1-sol or Sonnet 5.5
2.0M
One GPU's ceiling
~69M
Qwen3-Coder-30B-A3B on an RTX 5090 ($189)
Cheapest same-model API
93M (at the limit)
gpt-6-luna
60M
Claude Haiku 4.5
6.0M
gpt-6.1-sol or Sonnet 5.5
3.0M
One GPU's ceiling
~132M
gpt-oss-120b on an RTX PRO 6000 ($357)
Cheapest same-model API
313M (never)
gpt-6-luna
112M (never)
Claude Haiku 4.5
11.2M
gpt-6.1-sol or Sonnet 5.5
5.6M
One GPU's ceiling
~99M
Llama 3.3 70B on an RTX PRO 6000 ($357)
Cheapest same-model API
130M (never)
gpt-6-luna
112M (never)
Claude Haiku 4.5
11.2M
gpt-6.1-sol or Sonnet 5.5
5.6M
One GPU's ceiling
~30M
Cheapest same-model API
gpt-6-luna
Claude Haiku 4.5
gpt-6.1-sol or Sonnet 5.5
One GPU's ceiling
Llama 3.1 8B on an RTX 4090 ($128 a month)
262M (never)
40M
4.0M
2.0M
~69M
Qwen3-Coder-30B-A3B on an RTX 5090 ($189)
93M (at the limit)
60M
6.0M
3.0M
~132M
gpt-oss-120b on an RTX PRO 6000 ($357)
313M (never)
112M (never)
11.2M
5.6M
~99M
Llama 3.3 70B on an RTX PRO 6000 ($357)
130M (never)
112M (never)
11.2M
5.6M
~30M
The formula is the GPU's monthly cost divided by what the API charges for the same tokens:
break-even tokens a workday = (GPU $/hour × 176 hours + $68 admin) ÷ (22 workdays × blended API price per token)
The blended price assumes eight input tokens for every output token, which is typical when you send documents and get short answers back: (8 × input price + output price) ÷ 9. For Llama 3.1 8B on DeepInfra that's (8 × $0.02 + $0.04) ÷ 9 = $0.022 per million tokens, so an RTX 4090 at $128 a month needs 128 ÷ (22 × 0.022) = 262 million tokens a workday. The ceiling column is what one card can process in an 8-hour day if it's perfectly busy, from the benchmarks below.
Three things jump out:
The cheapest host sets the bar, and it's low. DeepInfra sells Llama 3.1 8B at $0.02 in and $0.04 out per million tokens. No rented card gets near that per token at a volume it can serve.
gpt-6-luna is an open model's toughest rival. At $0.10 in and $0.50 out, it's cheaper than renting a GPU for gpt-oss-120b or Llama 3.3 70B at any volume one card can handle.
Frontier models are where self-hosting wins. At $2 in and $10 out, gpt-6.1-sol and Claude Sonnet 5.5 cost enough that a $357-a-month GPU pays off from about 5.6 million tokens a workday. That's a quality decision, not just a hosting one: you're swapping a frontier model for an open one.
Which GPU each model needs, and what it rents for
Each model needs a card with enough memory for its weights plus the working memory for your context. The smallest card that fits a good-quality version runs from $0.34 an hour for Llama 3.1 8B to $1.64 an hour for gpt-oss-120b and Llama 3.3 70B, on Runpod on 2 October 2026.
Memory at 16k tokens of context. Runpod on-demand rates from aliteq's tracker, captured 2 Oct 2026, 19:21 UTC. · aliteq research
Llama 3.1 8B
Memory needed (16k context)
10.7 GB at 8-bit, 17.7 GB unquantized
Card we priced
RTX 4090, 24 GB
Runpod, per hour
$0.34
Vast.ai median, per hour
$0.50
Qwen3-Coder-30B-A3B
Memory needed (16k context)
25.6 GB at Q6_K, 19.5 GB at Q4_K_M
Card we priced
RTX 5090, 32 GB
Runpod, per hour
$0.69
Vast.ai median, per hour
$0.61
gpt-oss-120b
Memory needed (16k context)
67.9 GB at Q4_K_M, 91.1 GB at Q6_K
Card we priced
RTX PRO 6000 Max-Q, 96 GB
Runpod, per hour
$1.64
Vast.ai median, per hour
$1.60
Llama 3.3 70B
Memory needed (16k context)
45.6 GB at Q4_K_M, 75.6 GB at 8-bit
Card we priced
RTX PRO 6000 Max-Q, 96 GB
Runpod, per hour
$1.64
Vast.ai median, per hour
$1.60
Memory needed (16k context)
Card we priced
Runpod, per hour
Vast.ai median, per hour
Llama 3.1 8B
10.7 GB at 8-bit, 17.7 GB unquantized
RTX 4090, 24 GB
$0.34
$0.50
Qwen3-Coder-30B-A3B
25.6 GB at Q6_K, 19.5 GB at Q4_K_M
RTX 5090, 32 GB
$0.69
$0.61
gpt-oss-120b
67.9 GB at Q4_K_M, 91.1 GB at Q6_K
RTX PRO 6000 Max-Q, 96 GB
$1.64
$1.60
Llama 3.3 70B
45.6 GB at Q4_K_M, 75.6 GB at 8-bit
RTX PRO 6000 Max-Q, 96 GB
$1.64
$1.60
The memory figures for gpt-oss-120b and Qwen3-Coder-30B come straight from our cost to run pages. The two Llama pages can't show working memory yet because Meta's repository is licence-gated, so we applied the same formula to the model shapes published in public copies of Meta's config files. Llama 3.3 70B also fits a 48 GB card such as an L40S at 4-bit, with little room to spare. Every card that fits each model, with live prices, is on the model pages: gpt-oss-120b, Qwen3-Coder-30B, Llama 3.1 8B and Llama 3.3 70B.
We used Runpod's price because it's one fixed price per card type. Vast.ai is a marketplace of individual hosts, so we show its median listing, not its cheapest one, which moves hour to hour. Live rates for every card are on our GPU price table, and the best GPU for each model pages rank cards by fit and price.
Referral link
Testing an open model before you commit?
An RTX 4090 that runs Llama 3.1 8B was listed at $0.34 an hour on Runpod on 2 Oct 2026. Rent by the hour, stop the pod when you're done. Prices move.
Referral link — we may earn a commission at no cost to you. Prices on our compare page are the provider's live figures, cheapest first; this never changes the ranking.
How many tokens one GPU can actually serve
One GPU serving many users at once processes somewhere between about 1,000 and 4,600 tokens a second for these models, according to published benchmarks. That puts its ceiling at roughly 30 million to 130 million tokens in an 8-hour day, and you'll use only part of it.
Llama 3.1 8B, RTX 4090
Published throughput
~2,400 tokens/s
Ceiling, 8-hour day
~69M
Source
Spheron, vLLM, approximate
Qwen3-Coder-30B-A3B, RTX 5090
Published throughput
4,570 tokens/s
Ceiling, 8-hour day
~132M
Source
CloudRift, vLLM, 400 requests at once
gpt-oss-120b, RTX PRO 6000
Published throughput
~3,450 tokens/s
Ceiling, 8-hour day
~99M
Source
Derived from llama.cpp and CloudRift figures
Llama 3.3 70B, RTX PRO 6000
Published throughput
1,031 tokens/s
Ceiling, 8-hour day
~30M
Source
CloudRift, vLLM, 400 requests at once
Published throughput
Ceiling, 8-hour day
Source
Llama 3.1 8B, RTX 4090
~2,400 tokens/s
~69M
Spheron, vLLM, approximate
Qwen3-Coder-30B-A3B, RTX 5090
4,570 tokens/s
~132M
CloudRift, vLLM, 400 requests at once
gpt-oss-120b, RTX PRO 6000
~3,450 tokens/s
~99M
Derived from llama.cpp and CloudRift figures
Llama 3.3 70B, RTX PRO 6000
1,031 tokens/s
~30M
CloudRift, vLLM, 400 requests at once
These are ceilings measured with the card fully loaded. CloudRift's runs sent 400 requests at once, 1,000 tokens in and 1,000 out, and counted both directions; at that load each user waits a long time for the first word. Real traffic comes in peaks, so we treat 50% to 70% of the ceiling as usable. The two benchmark vendors rent GPUs, and Spheron's figure is a rounded one; both are published by companies that sell the self-hosting side.
One more caveat on the big card. The benchmarks used full-power RTX PRO 6000s, and the cheapest one on Runpod is the Max-Q, a 300-watt version that may be slower. Runpod's full-power card was $1.69 an hour, which moves the break-even figures up by about 2.5%.
Same model, different host: the price gap is huge
The same open model can cost ten times more on one API than on another, so who you compare against decides whether self-hosting looks smart. For Llama 3.3 70B, self-hosting beats Together's price at 15.6 million tokens a workday but never beats DeepInfra's.
Per million tokens, input / output, read 2 October 2026
Llama 3.1 8B
Cheapest list price we found
DeepInfra $0.02 / $0.04
Other hosts
Groq: contact sales
Qwen3-Coder-30B-A3B
Cheapest list price we found
Novita $0.07 / $0.27
Other hosts
Not listed on DeepInfra, Together or Groq
gpt-oss-120b
Cheapest list price we found
DeepInfra $0.037 / $0.17
Other hosts
Together, Groq, Fireworks $0.15 / $0.60
Llama 3.3 70B
Cheapest list price we found
DeepInfra $0.10 / $0.32
Other hosts
Together $1.04 / $1.04; Groq: contact sales
gpt-6-luna (closed)
Cheapest list price we found
OpenAI $0.10 / $0.50
Other hosts
Batch $0.05 / $0.25
Claude Haiku 4.5 (closed)
Cheapest list price we found
Anthropic $1 / $5
Other hosts
Batch half price
gpt-6.1-sol and Claude Sonnet 5.5 (closed)
Cheapest list price we found
$2 / $10
Other hosts
Batch half price
Cheapest list price we found
Other hosts
Llama 3.1 8B
DeepInfra $0.02 / $0.04
Groq: contact sales
Qwen3-Coder-30B-A3B
Novita $0.07 / $0.27
Not listed on DeepInfra, Together or Groq
gpt-oss-120b
DeepInfra $0.037 / $0.17
Together, Groq, Fireworks $0.15 / $0.60
Llama 3.3 70B
DeepInfra $0.10 / $0.32
Together $1.04 / $1.04; Groq: contact sales
gpt-6-luna (closed)
OpenAI $0.10 / $0.50
Batch $0.05 / $0.25
Claude Haiku 4.5 (closed)
Anthropic $1 / $5
Batch half price
gpt-6.1-sol and Claude Sonnet 5.5 (closed)
$2 / $10
Batch half price
Two details matter when you read that table:
The cheap Llama endpoints are quantized. DeepInfra's Turbo versions of Llama run at 8-bit (FP8), and the self-hosted figures in this article assume 4-bit to 8-bit too. Neither is the full-precision model, and quality differences are small but not zero.
Claude's newer models count more tokens. Anthropic says Claude 4.7 and later models produce about 30% more tokens for the same text, so Sonnet 5.5 costs more per word than its per-token price suggests. Haiku 4.5 uses the older tokenizer. Measured in words rather than tokens, self-hosting breaks even sooner against Sonnet 5.5 than the table shows.
Groq also lists Llama 3.1 8B and Llama 3.3 70B as enterprise models with no public price, and Fireworks prices models it doesn't list individually by size. We only used prices printed for the exact model.
What it costs at three volumes
At 1 million or 10 million tokens a workday, every API in this article costs less than renting a GPU for gpt-oss-120b. At 50 million, the rented card undercuts Claude and gpt-6.1-sol, but still not gpt-6-luna or any API for gpt-oss-120b itself.
22 workdays, eight input tokens per output token, no caching, list prices on 2 Oct 2026. · aliteq research
Monthly bill (22 workdays, 8:1 input to output, no caching)
gpt-oss-120b, DeepInfra
1M tokens a workday
$1.14
10M
$11
50M
$57
gpt-oss-120b, Together, Groq or Fireworks
1M tokens a workday
$4.40
10M
$44
50M
$220
gpt-6-luna
1M tokens a workday
$3.18
10M
$32
50M
$159
Rent an RTX PRO 6000 for gpt-oss-120b, 8 hours a day
1M tokens a workday
$357
10M
$357
50M
$357
Claude Haiku 4.5
1M tokens a workday
$32
10M
$318
50M
$1,589
gpt-6.1-sol or Claude Sonnet 5.5
1M tokens a workday
$64
10M
$636
50M
$3,178
1M tokens a workday
10M
50M
gpt-oss-120b, DeepInfra
$1.14
$11
$57
gpt-oss-120b, Together, Groq or Fireworks
$4.40
$44
$220
gpt-6-luna
$3.18
$32
$159
Rent an RTX PRO 6000 for gpt-oss-120b, 8 hours a day
$357
$357
$357
Claude Haiku 4.5
$32
$318
$1,589
gpt-6.1-sol or Claude Sonnet 5.5
$64
$636
$3,178
How big is 10 million tokens a day? A 25-person team asking an internal assistant 20 questions a day each, with about 4,000 tokens of documents and question going in and 500 coming back, uses about 2.25 million. So 1 million a workday is light team use, and 50 million is a product with real customers.
For small models, the fixed cost isn't the GPU. Our admin line is one hour a month at $68, the loaded hourly cost of a US systems administrator from the Bureau of Labor Statistics median wage. That's more than the $60 of RTX 4090 time a month.
The assumptions, and what they leave out
Every number above rests on choices we made. Change them and the break-even moves, but the pattern holds: the cheapest same-model API is hard to beat, and frontier prices are easy to beat.
Hours. Eight hours a day, 22 workdays, with the pod stopped at night: 176 GPU-hours a month. Run it around the clock and the GPU costs about 3.5 times as much; break-even against Sonnet 5.5 rises to 14.4 million tokens a day for gpt-oss-120b, and against the cheapest same-model API it's still out of reach for all four models except Qwen3-Coder-30B, which would need 204 million a day, about half its ceiling.
Token mix. Eight input tokens per output token. Chat with long answers has more output, which makes APIs dearer and helps the GPU a little.
Admin. One hour a month. A first deployment of vLLM or another serving stack takes far longer than that; we didn't price setup time.
Left out on purpose. Storage, data transfer, idle time while a model loads, a second GPU for failover, and prompt caching and batch discounts on the API side. All but the last make self-hosting more expensive.
Throughput. Published benchmarks from GPU sellers, not our tests, at full load. Your traffic pattern decides how much of that you get.
If you're thinking about buying the hardware instead of renting it, the three-year math is different. Our calculator compares an owned AI server with renting the same card:
$1.64/h × 1 GPU for the same hours · Vast median: $9,753
Break-even
18.1 h/day
Hours a day, every day for 3 years, before owning beats renting the same card.
For comparison: 25 staff using an internal document chat through an API costs about $356 (hosted open model) to $5,148 (Claude Sonnet 5 / gpt-6-sol) over the same three years, on our stated token assumptions.
Hardware: Bizon, Puget, NVIDIA and GMKtec pages; rental: aliteq GPU tracker (Runpod on-demand, Vast 30-day median); power: EIA, Eurostat (excl. VAT and recoverable taxes), DESNZ; admin wage: BLS. All checked 27 Sep 2026. Admin time, resale, usage and the 4-GPU price scaling are our assumptions. Power excludes cooling; colocation not included. An estimate, not a quote.
Self-hosting pays when it replaces a frontier model at a few million tokens a day, or when the reason isn't price. Here's how to decide.
Test whether an open model is good enough for your task. If it isn't, there's no break-even to calculate.
If it is, price that exact model on the cheapest host first. DeepInfra, Novita, Together, Groq and Fireworks all list prices per model.
Estimate your tokens a workday. Under about 2 million, any API is cheaper than a dedicated GPU.
Compare with the break-even table. Against a frontier model, a rented GPU pays from about 2 to 11 million tokens a workday.
Self-host anyway if prompts must stay on hardware you control, you run a fine-tuned model no host offers, or rate limits block you.
If keeping data inside the EU is the reason, the routes and prices are different; we cover them in private LLM cost for EU companies. More cost guides for running AI in a business are on the AI Automation hub.
Quick answers
Is it cheaper to self-host an LLM or use an API?
Usually the API. One rented GPU running Llama 3.1 8B, gpt-oss-120b or Llama 3.3 70B never breaks even against the cheapest API for the same model at a volume the card can serve. Self-hosting wins against frontier models such as gpt-6.1-sol or Claude Sonnet 5.5, from about 2 to 6 million tokens a workday.
How many tokens a day do I need before self-hosting pays off?
Against gpt-6.1-sol or Claude Sonnet 5.5, about 2 million tokens a workday for Llama 3.1 8B and 5.6 million for gpt-oss-120b or Llama 3.3 70B, on one rented GPU for 8 hours a day. Against Claude Haiku 4.5 it's 4 to 11 million. Against the same open model on DeepInfra, it's 130 to 313 million, beyond one GPU.
What GPU do I need to run gpt-oss-120b?
An 80 to 96 GB card. gpt-oss-120b needs about 67.9 GB at Q4_K_M with 16k tokens of context, so it fits an RTX PRO 6000 (96 GB), which rented for $1.64 an hour on Runpod on 2 October 2026, or an 80 GB H100.
What is the cheapest API for gpt-oss-120b?
Of the vendor pages we read on 2 October 2026, DeepInfra at $0.037 per million input tokens and $0.17 per million output tokens. Together, Groq and Fireworks list it at $0.15 and $0.60.
How many tokens can one GPU serve in a day?
Published benchmarks put one card's ceiling at roughly 30 million tokens in an 8-hour day for Llama 3.3 70B on an RTX PRO 6000, up to about 132 million for Qwen3-Coder-30B on an RTX 5090, at full load. Plan on using 50% to 70% of that.
Does running the GPU 24/7 change the break-even?
It raises it. Around the clock, a rented RTX PRO 6000 costs about $1,265 a month, so gpt-oss-120b breaks even against Claude Sonnet 5.5 at about 14.4 million tokens a day instead of 5.6 million a workday. Against the cheapest same-model API it still doesn't break even.
Use this in your own page
Teaching this? Paste the live version into your course, blog or answer. Free, no sign-up; the credit line links back here.