aliteq.

Self-Hosted LLM vs API Break-Even: Tokens per Day by Model (Live GPU Prices)

Renting a GPU for an open model looks cheaper than paying per token until you price the same model on the cheapest API. We did it for four open models, from Llama 3.1 8B to Llama 3.3 70B, with live GPU rates from our own tracker and list prices read on 2 October 2026. Against the same model's API, one rented GPU almost never wins. Against a frontier model, it wins from about 2 to 6 million tokens a day.

TensorUpdated 1h ago12 min readWeb story
Hand-drawn editorial illustration of a level seesaw balancing a graphics card on one end against a tall stack of lime-green coins on the other

This post contains affiliate links. If you buy through them, Aliteq may earn a commission — at no extra cost to you. Prices verified at publish time.

Share

"Should we run our own model?" usually gets answered with a GPU's hourly price and a frontier model's per-token price. That comparison flatters the GPU. The fair test is the same open model on someone else's GPU, billed per token, and those hosts are cheap.

So we priced both sides for four open models of different sizes. The GPU side uses aliteq's own tracker, which has logged Runpod's and Vast's prices every hour since 21 July 2026; this snapshot was captured at 19:21 UTC on 2 October 2026. For the API side we read each vendor's pricing page on 2 October 2026. The throughput figures are other people's published benchmarks. We didn't deploy or benchmark any of these models ourselves.

The break-even, model by model

One rented GPU beats the cheapest same-model API only if you push tens of millions of tokens through it every workday, and for three of our four models that's more than the card can serve. Against closed frontier models, break-even drops to between 2 and 11 million tokens a workday.

Chart of break-even tokens per workday for running four open models on one rented GPU in business hours. Against the cheapest API for the same model: Llama 3.1 8B 262 million, Qwen3-Coder-30B 93 million, gpt-oss-120b 313 million, Llama 3.3 70B 130 million. Against Claude Haiku 4.5: 4.0, 6.0, 11.2 and 11.2 million. Against gpt-6.1-sol or Claude Sonnet 5.5: 2.0, 3.0, 5.6 and 5.6 million.
Millions of tokens a workday at which one rented GPU in business hours costs the same as the API. Above one GPU's ceiling, it never breaks even. · aliteq research

Break-even, millions of tokens per workday (one rented GPU, 8 hours a day)

Llama 3.1 8B on an RTX 4090 ($128 a month)

Cheapest same-model API
262M (never)
gpt-6-luna
40M
Claude Haiku 4.5
4.0M
gpt-6.1-sol or Sonnet 5.5
2.0M
One GPU's ceiling
~69M

Qwen3-Coder-30B-A3B on an RTX 5090 ($189)

Cheapest same-model API
93M (at the limit)
gpt-6-luna
60M
Claude Haiku 4.5
6.0M
gpt-6.1-sol or Sonnet 5.5
3.0M
One GPU's ceiling
~132M

gpt-oss-120b on an RTX PRO 6000 ($357)

Cheapest same-model API
313M (never)
gpt-6-luna
112M (never)
Claude Haiku 4.5
11.2M
gpt-6.1-sol or Sonnet 5.5
5.6M
One GPU's ceiling
~99M

Llama 3.3 70B on an RTX PRO 6000 ($357)

Cheapest same-model API
130M (never)
gpt-6-luna
112M (never)
Claude Haiku 4.5
11.2M
gpt-6.1-sol or Sonnet 5.5
5.6M
One GPU's ceiling
~30M

The formula is the GPU's monthly cost divided by what the API charges for the same tokens:

break-even tokens a workday = (GPU $/hour × 176 hours + $68 admin) ÷ (22 workdays × blended API price per token)

The blended price assumes eight input tokens for every output token, which is typical when you send documents and get short answers back: (8 × input price + output price) ÷ 9. For Llama 3.1 8B on DeepInfra that's (8 × $0.02 + $0.04) ÷ 9 = $0.022 per million tokens, so an RTX 4090 at $128 a month needs 128 ÷ (22 × 0.022) = 262 million tokens a workday. The ceiling column is what one card can process in an 8-hour day if it's perfectly busy, from the benchmarks below.

Three things jump out:

  • The cheapest host sets the bar, and it's low. DeepInfra sells Llama 3.1 8B at $0.02 in and $0.04 out per million tokens. No rented card gets near that per token at a volume it can serve.
  • gpt-6-luna is an open model's toughest rival. At $0.10 in and $0.50 out, it's cheaper than renting a GPU for gpt-oss-120b or Llama 3.3 70B at any volume one card can handle.
  • Frontier models are where self-hosting wins. At $2 in and $10 out, gpt-6.1-sol and Claude Sonnet 5.5 cost enough that a $357-a-month GPU pays off from about 5.6 million tokens a workday. That's a quality decision, not just a hosting one: you're swapping a frontier model for an open one.

Which GPU each model needs, and what it rents for

Each model needs a card with enough memory for its weights plus the working memory for your context. The smallest card that fits a good-quality version runs from $0.34 an hour for Llama 3.1 8B to $1.64 an hour for gpt-oss-120b and Llama 3.3 70B, on Runpod on 2 October 2026.

Table of four open models, the GPU each fits on and its rental price on 2 October 2026: Llama 3.1 8B needs 10.7 GB at 8-bit and fits an RTX 4090 at $0.34 an hour; Qwen3-Coder-30B needs 25.6 GB at Q6_K and fits an RTX 5090 at $0.69; gpt-oss-120b needs 67.9 GB at Q4_K_M and fits an RTX PRO 6000 at $1.64; Llama 3.3 70B needs 75.6 GB at 8-bit and fits an RTX PRO 6000 at $1.64.
Memory at 16k tokens of context. Runpod on-demand rates from aliteq's tracker, captured 2 Oct 2026, 19:21 UTC. · aliteq research

Llama 3.1 8B

Memory needed (16k context)
10.7 GB at 8-bit, 17.7 GB unquantized
Card we priced
RTX 4090, 24 GB
Runpod, per hour
$0.34
Vast.ai median, per hour
$0.50

Qwen3-Coder-30B-A3B

Memory needed (16k context)
25.6 GB at Q6_K, 19.5 GB at Q4_K_M
Card we priced
RTX 5090, 32 GB
Runpod, per hour
$0.69
Vast.ai median, per hour
$0.61

gpt-oss-120b

Memory needed (16k context)
67.9 GB at Q4_K_M, 91.1 GB at Q6_K
Card we priced
RTX PRO 6000 Max-Q, 96 GB
Runpod, per hour
$1.64
Vast.ai median, per hour
$1.60

Llama 3.3 70B

Memory needed (16k context)
45.6 GB at Q4_K_M, 75.6 GB at 8-bit
Card we priced
RTX PRO 6000 Max-Q, 96 GB
Runpod, per hour
$1.64
Vast.ai median, per hour
$1.60

The memory figures for gpt-oss-120b and Qwen3-Coder-30B come straight from our cost to run pages. The two Llama pages can't show working memory yet because Meta's repository is licence-gated, so we applied the same formula to the model shapes published in public copies of Meta's config files. Llama 3.3 70B also fits a 48 GB card such as an L40S at 4-bit, with little room to spare. Every card that fits each model, with live prices, is on the model pages: gpt-oss-120b, Qwen3-Coder-30B, Llama 3.1 8B and Llama 3.3 70B.

We used Runpod's price because it's one fixed price per card type. Vast.ai is a marketplace of individual hosts, so we show its median listing, not its cheapest one, which moves hour to hour. Live rates for every card are on our GPU price table, and the best GPU for each model pages rank cards by fit and price.

RunpodReferral link

Testing an open model before you commit?

An RTX 4090 that runs Llama 3.1 8B was listed at $0.34 an hour on Runpod on 2 Oct 2026. Rent by the hour, stop the pod when you're done. Prices move.

Referral link — we may earn a commission at no cost to you. Prices on our compare page are the provider's live figures, cheapest first; this never changes the ranking.

How many tokens one GPU can actually serve

One GPU serving many users at once processes somewhere between about 1,000 and 4,600 tokens a second for these models, according to published benchmarks. That puts its ceiling at roughly 30 million to 130 million tokens in an 8-hour day, and you'll use only part of it.

Llama 3.1 8B, RTX 4090

Published throughput
~2,400 tokens/s
Ceiling, 8-hour day
~69M
Source
Spheron, vLLM, approximate

Qwen3-Coder-30B-A3B, RTX 5090

Published throughput
4,570 tokens/s
Ceiling, 8-hour day
~132M
Source
CloudRift, vLLM, 400 requests at once

gpt-oss-120b, RTX PRO 6000

Published throughput
~3,450 tokens/s
Ceiling, 8-hour day
~99M
Source
Derived from llama.cpp and CloudRift figures

Llama 3.3 70B, RTX PRO 6000

Published throughput
1,031 tokens/s
Ceiling, 8-hour day
~30M
Source
CloudRift, vLLM, 400 requests at once

These are ceilings measured with the card fully loaded. CloudRift's runs sent 400 requests at once, 1,000 tokens in and 1,000 out, and counted both directions; at that load each user waits a long time for the first word. Real traffic comes in peaks, so we treat 50% to 70% of the ceiling as usable. The two benchmark vendors rent GPUs, and Spheron's figure is a rounded one; both are published by companies that sell the self-hosting side.

One more caveat on the big card. The benchmarks used full-power RTX PRO 6000s, and the cheapest one on Runpod is the Max-Q, a 300-watt version that may be slower. Runpod's full-power card was $1.69 an hour, which moves the break-even figures up by about 2.5%.

Same model, different host: the price gap is huge

The same open model can cost ten times more on one API than on another, so who you compare against decides whether self-hosting looks smart. For Llama 3.3 70B, self-hosting beats Together's price at 15.6 million tokens a workday but never beats DeepInfra's.

Per million tokens, input / output, read 2 October 2026

Llama 3.1 8B

Cheapest list price we found
DeepInfra $0.02 / $0.04
Other hosts
Groq: contact sales

Qwen3-Coder-30B-A3B

Cheapest list price we found
Novita $0.07 / $0.27
Other hosts
Not listed on DeepInfra, Together or Groq

gpt-oss-120b

Cheapest list price we found
DeepInfra $0.037 / $0.17
Other hosts
Together, Groq, Fireworks $0.15 / $0.60

Llama 3.3 70B

Cheapest list price we found
DeepInfra $0.10 / $0.32
Other hosts
Together $1.04 / $1.04; Groq: contact sales

gpt-6-luna (closed)

Cheapest list price we found
OpenAI $0.10 / $0.50
Other hosts
Batch $0.05 / $0.25

Claude Haiku 4.5 (closed)

Cheapest list price we found
Anthropic $1 / $5
Other hosts
Batch half price

gpt-6.1-sol and Claude Sonnet 5.5 (closed)

Cheapest list price we found
$2 / $10
Other hosts
Batch half price

Two details matter when you read that table:

  • The cheap Llama endpoints are quantized. DeepInfra's Turbo versions of Llama run at 8-bit (FP8), and the self-hosted figures in this article assume 4-bit to 8-bit too. Neither is the full-precision model, and quality differences are small but not zero.
  • Claude's newer models count more tokens. Anthropic says Claude 4.7 and later models produce about 30% more tokens for the same text, so Sonnet 5.5 costs more per word than its per-token price suggests. Haiku 4.5 uses the older tokenizer. Measured in words rather than tokens, self-hosting breaks even sooner against Sonnet 5.5 than the table shows.

Groq also lists Llama 3.1 8B and Llama 3.3 70B as enterprise models with no public price, and Fireworks prices models it doesn't list individually by size. We only used prices printed for the exact model.

What it costs at three volumes

At 1 million or 10 million tokens a workday, every API in this article costs less than renting a GPU for gpt-oss-120b. At 50 million, the rented card undercuts Claude and gpt-6.1-sol, but still not gpt-6-luna or any API for gpt-oss-120b itself.

Table of monthly cost at 1, 10 and 50 million tokens a workday: gpt-oss-120b on DeepInfra $1.14, $11 and $57; on Together, Groq or Fireworks $4.40, $44 and $220; gpt-6-luna $3.18, $32 and $159; renting an RTX PRO 6000 for gpt-oss-120b in business hours $357 at every volume; Claude Haiku 4.5 $32, $318 and $1,589; gpt-6.1-sol or Claude Sonnet 5.5 $64, $636 and $3,178.
22 workdays, eight input tokens per output token, no caching, list prices on 2 Oct 2026. · aliteq research

Monthly bill (22 workdays, 8:1 input to output, no caching)

gpt-oss-120b, DeepInfra

1M tokens a workday
$1.14
10M
$11
50M
$57

gpt-oss-120b, Together, Groq or Fireworks

1M tokens a workday
$4.40
10M
$44
50M
$220

gpt-6-luna

1M tokens a workday
$3.18
10M
$32
50M
$159

Rent an RTX PRO 6000 for gpt-oss-120b, 8 hours a day

1M tokens a workday
$357
10M
$357
50M
$357

Claude Haiku 4.5

1M tokens a workday
$32
10M
$318
50M
$1,589

gpt-6.1-sol or Claude Sonnet 5.5

1M tokens a workday
$64
10M
$636
50M
$3,178

How big is 10 million tokens a day? A 25-person team asking an internal assistant 20 questions a day each, with about 4,000 tokens of documents and question going in and 500 coming back, uses about 2.25 million. So 1 million a workday is light team use, and 50 million is a product with real customers.

For small models, the fixed cost isn't the GPU. Our admin line is one hour a month at $68, the loaded hourly cost of a US systems administrator from the Bureau of Labor Statistics median wage. That's more than the $60 of RTX 4090 time a month.

The assumptions, and what they leave out

Every number above rests on choices we made. Change them and the break-even moves, but the pattern holds: the cheapest same-model API is hard to beat, and frontier prices are easy to beat.

  • Hours. Eight hours a day, 22 workdays, with the pod stopped at night: 176 GPU-hours a month. Run it around the clock and the GPU costs about 3.5 times as much; break-even against Sonnet 5.5 rises to 14.4 million tokens a day for gpt-oss-120b, and against the cheapest same-model API it's still out of reach for all four models except Qwen3-Coder-30B, which would need 204 million a day, about half its ceiling.
  • Token mix. Eight input tokens per output token. Chat with long answers has more output, which makes APIs dearer and helps the GPU a little.
  • Admin. One hour a month. A first deployment of vLLM or another serving stack takes far longer than that; we didn't price setup time.
  • Left out on purpose. Storage, data transfer, idle time while a model loads, a second GPU for failover, and prompt caching and batch discounts on the API side. All but the last make self-hosting more expensive.
  • Throughput. Published benchmarks from GPU sellers, not our tests, at full load. Your traffic pattern decides how much of that you get.

If you're thinking about buying the hardware instead of renting it, the three-year math is different. Our calculator compares an owned AI server with renting the same card:

3-year cost: own vs rent

Own it

$32,029

box $26,979 + power $654 + admin $9,792 − resale $5,396

Rent it on Runpod

$10,391

$1.64/h × 1 GPU for the same hours · Vast median: $9,753

Break-even

18.1 h/day

Hours a day, every day for 3 years, before owning beats renting the same card.

For comparison: 25 staff using an internal document chat through an API costs about $356 (hosted open model) to $5,148 (Claude Sonnet 5 / gpt-6-sol) over the same three years, on our stated token assumptions.

Hardware: Bizon, Puget, NVIDIA and GMKtec pages; rental: aliteq GPU tracker (Runpod on-demand, Vast 30-day median); power: EIA, Eurostat (excl. VAT and recoverable taxes), DESNZ; admin wage: BLS. All checked 27 Sep 2026. Admin time, resale, usage and the 4-GPU price scaling are our assumptions. Power excludes cooling; colocation not included. An estimate, not a quote.

The full write-up is in AI server vs cloud GPU: the 3-year cost.

When self-hosting does make sense

Self-hosting pays when it replaces a frontier model at a few million tokens a day, or when the reason isn't price. Here's how to decide.

Test whether an open model is good enough for your task. If it isn't, there's no break-even to calculate.

If it is, price that exact model on the cheapest host first. DeepInfra, Novita, Together, Groq and Fireworks all list prices per model.

Estimate your tokens a workday. Under about 2 million, any API is cheaper than a dedicated GPU.

Compare with the break-even table. Against a frontier model, a rented GPU pays from about 2 to 11 million tokens a workday.

Self-host anyway if prompts must stay on hardware you control, you run a fine-tuned model no host offers, or rate limits block you.

If keeping data inside the EU is the reason, the routes and prices are different; we cover them in private LLM cost for EU companies. More cost guides for running AI in a business are on the AI Automation hub.

Quick answers

Is it cheaper to self-host an LLM or use an API?
Usually the API. One rented GPU running Llama 3.1 8B, gpt-oss-120b or Llama 3.3 70B never breaks even against the cheapest API for the same model at a volume the card can serve. Self-hosting wins against frontier models such as gpt-6.1-sol or Claude Sonnet 5.5, from about 2 to 6 million tokens a workday.
How many tokens a day do I need before self-hosting pays off?
Against gpt-6.1-sol or Claude Sonnet 5.5, about 2 million tokens a workday for Llama 3.1 8B and 5.6 million for gpt-oss-120b or Llama 3.3 70B, on one rented GPU for 8 hours a day. Against Claude Haiku 4.5 it's 4 to 11 million. Against the same open model on DeepInfra, it's 130 to 313 million, beyond one GPU.
What GPU do I need to run gpt-oss-120b?
An 80 to 96 GB card. gpt-oss-120b needs about 67.9 GB at Q4_K_M with 16k tokens of context, so it fits an RTX PRO 6000 (96 GB), which rented for $1.64 an hour on Runpod on 2 October 2026, or an 80 GB H100.
What is the cheapest API for gpt-oss-120b?
Of the vendor pages we read on 2 October 2026, DeepInfra at $0.037 per million input tokens and $0.17 per million output tokens. Together, Groq and Fireworks list it at $0.15 and $0.60.
How many tokens can one GPU serve in a day?
Published benchmarks put one card's ceiling at roughly 30 million tokens in an 8-hour day for Llama 3.3 70B on an RTX PRO 6000, up to about 132 million for Qwen3-Coder-30B on an RTX 5090, at full load. Plan on using 50% to 70% of that.
Does running the GPU 24/7 change the break-even?
It raises it. Around the clock, a rented RTX PRO 6000 costs about $1,265 a month, so gpt-oss-120b breaks even against Claude Sonnet 5.5 at about 14.4 million tokens a day instead of 5.6 million a workday. Against the cheapest same-model API it still doesn't break even.

Use this in your own page

Teaching this? Paste the live version into your course, blog or answer. Free, no sign-up; the credit line links back here.

Embed
Cite

Found this useful? Share it

Share
Tensor

Local AI & Automation Editor

Tensor

I'm US-based, I run more models at home than I'll admit to, and I've quantized more than I've finished reading about. I write about running AI on your own hardware and, lately, about what it costs a company to do the same — tokens per day, GPUs per month, and the GDPR questions nobody's sales deck answers.

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading