aliteq.

Cloud GPU vs Your Own AI Server: The 3-Year Cost for a Small Company (Calculator)

Should a 10–50 person company buy a GPU server, rent one by the hour, or just pay per token? The three-year bill for each, from live rental prices, today's hardware prices and government power statistics. The answer for most small companies is not the one GPU vendors sell.

0xLaxUpdated 56m ago13 min readWeb story
Hand-drawn illustration of a small office with a warm GPU tower under a desk and a cloud with a price tag outside the window
Share

I spent five years building, repairing and selling PCs at Aliteq's shop in Kathmandu, from 2019 to 2024. These days I build AI apps from Denmark. So when a small company asks whether it should buy its own AI server, I understand why the idea is attractive: you pay once, the machine is yours, and your data never leaves the office.

The numbers usually say otherwise, and the pages that rank for this question won't tell you, because every one of them sells one side of the answer. Lenovo's TCO whitepaper sells servers. Runpod's and Spheron's sell rented GPUs. Lenovo's smallest configuration starts at two RTX PRO 6000 cards for about $68,000, and none of them model a company with one box, a handful of users, or an API bill. aliteq sells neither hardware nor cloud. We do have referral links with Vast and Runpod, disclosed wherever they appear, and they don't change the maths here.

What we do have is our own data. aliteq has logged Runpod's and Vast's GPU prices every hour since 21 July 2026, so the rental numbers below are ours, not a vendor's. The hardware prices were read on retailer and builder pages on 27 September 2026, and the power prices come from government statistics.

The answer in one chart

For a typical small company, the API is the cheapest option by a wide margin, renting a GPU by the hour comes next, and owning comes last. Over three years, a 25-person internal document chat costs $356 to $5,148 through an API, about $10,391 on a rented RTX PRO 6000 in business hours, and about $32,029 on an owned one.

Bar chart of three-year cost for a 25-person company's internal document chat: hosted gpt-oss-120b API $356, Claude Sonnet 5 or gpt-6-sol API $5,148, renting an RTX PRO 6000 on Runpod in business hours $10,391, owning a Strix Halo mini-PC $12,714, 25 chat seats $18,000, owning a 1× RTX PRO 6000 workstation $32,029.
The same workload, priced six ways over three years. Assumptions are listed below the chart and in the calculator. · aliteq research

How much AI work does a small company actually do?

Far less than GPU vendors assume. Their break-even maths assumes a box busy 60–70% of the time. By my estimate, a 25-person document chat needs about four hours of real GPU work a month on one RTX PRO 6000, and a month of batch document processing about two. A coding assistant for ten developers is the exception.

No vendor publishes how many tokens a chat question or a document page uses, so these three workloads are my assumptions, stated in full:

Three workloads (my assumptions)

Document chat, 25 staff

Assumption
20 questions a workday each; 4,000 tokens in, 500 out per question
Tokens per month
44M in + 5.5M out
GPU time a month (1× RTX PRO 6000)
~4 hours of work

Coding assistant, 10 developers

Assumption
20M input tokens per developer-day, 90% of them cached; 150k out
Tokens per month
4,400M in (3,960M cached) + 33M out
GPU time a month (1× RTX PRO 6000)
~30–140 hours

Document processing

Assumption
20,000 pages; 800 tokens in and 200 out per page
Tokens per month
16M in + 4M out
GPU time a month (1× RTX PRO 6000)
~2 hours

The coding profile is calibrated to the one real figure I could find: Anthropic's documentation reports that Claude Code costs enterprises about $13 per developer per active day on average. My assumption works out to about $9 a day on Sonnet 5 and $15 on Opus 5.5, which brackets it.

The GPU-time column uses published benchmarks, not tests of our own: figures in llama.cpp's gpt-oss guide show about 4,500 tokens a second of prompt processing for gpt-oss-120b on one RTX PRO 6000. The point isn't the exact hours. It's that a small company's chat and document work keeps a dedicated GPU idle nearly all day.

What the API and seats cost over three years

Per-token pricing makes the chat and document workloads almost free on open models and a few thousand dollars on frontier ones. Coding is where the bill grows: tens of thousands of dollars on API pricing, or a few thousand on per-seat coding plans.

3-year cost by model (list prices, 27 Sep 2026)

gpt-oss-120b, hosted ($0.15 in / $0.60 out per 1M)

Document chat, 25 staff
$356
Coding, 10 developers
$24,473
Documents (batch)
$173

gpt-6-luna ($0.10 / $0.50)

Document chat, 25 staff
$257
Coding, 10 developers
$3,604
Documents (batch)
$65

Claude Haiku 4.5 ($1 / $5)

Document chat, 25 staff
$2,574
Coding, 10 developers
$36,036
Documents (batch)
$648

Claude Sonnet 5 or gpt-6-sol ($2 / $10)

Document chat, 25 staff
$5,148
Coding, 10 developers
$72,072
Documents (batch)
$1,296

Claude Opus 5.5 ($4 / $20)

Document chat, 25 staff
$10,296
Coding, 10 developers
$115,632
Documents (batch)
$2,592

Seats change the picture for people rather than systems. Twenty-five ChatGPT Business or Claude Team seats cost $18,000 over three years at $20 a month billed annually, but they include the chat app, admin controls and SSO that a raw API doesn't. For ten developers, GitHub Copilot Business comes to $6,840 over three years and Claude Team standard seats, which include Claude Code, to $7,200.

What owning actually costs

A box's three-year bill is its price, plus power, plus someone's time, minus what you sell it for. For one RTX PRO 6000 workstation in business hours that's about $32,029: $26,979 for the machine, $654 of US power, $9,792 of admin time at four hours a month, less a 20% resale credit.

Own vs rent, 3 years (US power, business hours, 4 h/month admin, 20% resale)

DIY box, 1× RTX 5090 ($7,300)

Own, 3 years
$16,435
Rent the same card, business hours (Runpod)
$4,372
Rent 24/7 (Runpod)
$18,133

1× RTX PRO 6000 Max-Q ($26,979)

Own, 3 years
$32,029
Rent the same card, business hours (Runpod)
$10,391
Rent 24/7 (Runpod)
$43,099

2× RTX PRO 6000 Max-Q ($51,858)

Own, 3 years
$52,333
Rent the same card, business hours (Runpod)
$20,782
Rent 24/7 (Runpod)
$86,198

4× RTX PRO 6000 Max-Q (~$84,600)

Own, 3 years
$79,192
Rent the same card, business hours (Runpod)
$41,564
Rent 24/7 (Runpod)
$172,397

NVIDIA DGX Spark ($4,699)

Own, 3 years
$13,776
Rent the same card, business hours (Runpod)
not tracked
Rent 24/7 (Runpod)
not tracked

Strix Halo mini-PC ($3,500)

Own, 3 years
$12,714
Rent the same card, business hours (Runpod)
not tracked
Rent 24/7 (Runpod)
not tracked

Change any of it here. The defaults are the numbers above, and the admin time, resale value and hours are sliders because they're my assumptions:

3-year cost: own vs rent

Own it

$32,029

box $26,979 + power $654 + admin $9,792 − resale $5,396

Rent it on Runpod

$10,391

$1.64/h × 1 GPU for the same hours · Vast median: $9,753

Break-even

18.1 h/day

Hours a day, every day for 3 years, before owning beats renting the same card.

For comparison: 25 staff using an internal document chat through an API costs about $356 (hosted open model) to $5,148 (Claude Sonnet 5 / gpt-6-sol) over the same three years, on our stated token assumptions.

Hardware: Bizon, Puget, NVIDIA and GMKtec pages; rental: aliteq GPU tracker (Runpod on-demand, Vast 30-day median); power: EIA, Eurostat (excl. VAT and recoverable taxes), DESNZ; admin wage: BLS. All checked 27 Sep 2026. Admin time, resale, usage and the 4-GPU price scaling are our assumptions. Power excludes cooling; colocation not included. An estimate, not a quote.

Three things about that table surprise people:

  • Admin time is the biggest hidden line. Nobody measures how long one small inference box takes to look after, so I use four hours a month, below the 10–20 hours some guides assume for production serving. At a loaded US sysadmin cost of about $68 an hour (the BLS median wage of $47.66, divided by the 70% wages make up of total pay), that's $9,792 over three years. For a DGX Spark or a Strix Halo box, admin is 71–77% of the whole bill.
  • Power barely matters. It's 2–17% of the cost of owning in every case I modelled. One RTX PRO 6000 box costs about $654 in US power over three years in business hours, $1,027 in Denmark and $2,190 in the UK, using government statistics for business electricity.
  • Hardware is expensive right now. The RTX 5090's $1,999 launch price can't be paid anywhere I checked: partner cards list at $4,400–$4,900 and are sold out, and an in-stock one was $7,980 at Best Buy. An RTX PRO 6000 card is $11,739–$16,269 and backordered. The DGX Spark is $4,699 and the GMKtec Strix Halo box with 128 GB is $3,499.99.

On taxes, owning has one real advantage: in the US, IRS Publication 946 treats computers as five-year property, and for tax years beginning in 2026 the Section 179 deduction goes up to $2,560,000, so a company can often write off a server in the year it buys it. In Denmark, equipment depreciates at up to 25% a year on a declining balance. Ask your accountant; this isn't tax advice.

How many hours a day your server has to work

To beat renting the same card on Runpod, one RTX PRO 6000 box has to work about 18 hours a day, every day, for three years. A DIY RTX 5090 box needs 23.5 hours, which is never in practice. Only a four-GPU box gets close to a normal working day, at about 11 hours.

Range bars of break-even hours per day over three years versus renting on Runpod: DIY RTX 5090 box 8.8 to 23.5 hours, one RTX PRO 6000 Max-Q 12.5 to 18.1 hours, two 11.9 to 14.7 hours, four 9.7 to 11.1 hours.
Left end: admin time is free. Right end: four hours a month at $68 an hour. · aliteq research

If someone looks after the box for free, because it's their hobby or already their job, the numbers improve a lot: 8.8 hours a day for the RTX 5090 box and 12.5 for one RTX PRO 6000. That's why admin time, not the GPU price or the power bill, decides the single-box case. Be honest with yourself about who will patch it, update the drivers and answer "the AI is down" at 9 in the morning.

Where to rent, if you rent

Rent from a specialist GPU cloud, not a hyperscaler. One H100 costs $6.88 an hour on AWS against $2.69 on Runpod and a $2.18 median on Vast. For a small company's private open model, a 96 GB RTX PRO 6000 at about $1.64 an hour is usually enough.

Bar chart of one GPU-hour on demand: H100 on AWS $6.88, Runpod $2.69, Vast median $2.18; L40S on AWS $1.86, Runpod $0.79, Vast $0.80; RTX PRO 6000 on Runpod $1.64, Vast $1.72; RTX 5090 on Runpod $0.69, Vast $0.60.
Snapshot from aliteq's GPU price tracker, 27 September 2026; AWS on-demand price feed. · aliteq research

Our tracker also contradicts a story you may have seen, that GPU rental prices have doubled. Since we started logging on 21 July 2026, Runpod's list prices haven't changed at all, and Vast's median H100 price has fallen 24%. The one card that rose is the RTX PRO 6000, up 49% on Vast, which is consistent with its purchase price climbing. The live table updates every hour.

If you want a whole machine by the month in the EU instead, Hetzner's GEX131 with one RTX PRO 6000 Max-Q costs €1,197.30 a month plus €599 setup, which is about $49,833 over three years at today's exchange rate. That's more than owning the same card and more than renting it around the clock on Runpod.

When owning is the right call

Own the server when your data must not leave the building, or when a steady workload keeps the box busy most of every day. A coding assistant for a large team, round-the-clock batch jobs, or a four-GPU box that's used all day can justify it. A chat assistant for 25 people can't.

The data reason is real and often decisive. If a customer contract, a regulator or your own risk appetite rules out sending documents to a third party, a local model on your own hardware removes the question entirely. The Security & Compliance section covers what customers actually ask for, and SOC 2 for AI startups covers what your API vendors already cover.

If you do buy, plan the room before the box. Every watt a server draws becomes heat: about 3,412 BTU an hour per kilowatt, according to the EIA. Four 600-watt cards need more power than a normal office circuit is built for, so check with an electrician first. Colocation is the alternative, but cheap single-server deals include only 60–360 watts. FDC Servers' high-density offer, $999 a month with 4 kW included, is the kind of plan a four-GPU box needs.

What I'd do at a 10–50 person company

Start with an API or a handful of seats. If you need a private open model, rent the GPU by the hour until you can see a steady, all-day workload in your own usage logs. Buy hardware only when data control demands it or your logs prove the box would be busy.

That matches the plain answer in does my business need a GPU for AI: almost certainly not. The numbers above are why. If you do go the self-hosted route, running a local model as an API server and the per-model cost to run pages show which GPUs fit which models.

Quick answers

Is it cheaper to buy an AI server or rent cloud GPUs for a small company?
Usually renting, and often an API is cheaper still. One RTX PRO 6000 box costs about $32,029 over three years including admin and resale, against about $10,391 to rent the same card on Runpod for business hours. Owning only wins if the box works about 18 hours a day for three years.
How many users can one GPU server handle?
There's no fixed number. It depends on the model, prompt length and how many people ask at the same moment. For a small company's document chat, one RTX PRO 6000 has far more capacity than 25 people use; the real limit is usually how many ask at once, not the total.
How much does it cost to run an AI server in electricity?
Less than people expect. One RTX PRO 6000 workstation costs about $654 over three years in US business electricity during business hours, about $1,565 if it runs flat out around the clock. Power is 2–17% of the total cost of owning in every case we modelled.
Is self-hosting an LLM cheaper than using the OpenAI or Claude API?
For small companies, rarely. A 25-person document chat costs about $356 over three years on a hosted open model and $5,148 on Claude Sonnet 5 or gpt-6-sol, while owning a capable box costs over $12,000 once admin time is included. Self-hosting pays off for data control or very heavy, steady use.
What is a GPU server worth after three years?
In a normal market, a fraction of its price; I assume 20%. Today's shortage has pushed used prices unusually high: a used RTX 4090 sells for more than its launch price. Don't count on that lasting.
Should a small company use AWS for GPUs?
Not for simple GPU rental. AWS charges $6.88 an hour for one H100, against $2.69 on Runpod. Hyperscalers make sense when you need their other services, compliance paperwork or existing contracts.

The rest of the Cloud & IT Infrastructure section prices the other decisions a small company faces: managed IT, backups, device management and the Windows Server 2016 deadline.

Use this in your own page

Teaching this? Paste the live version into your course, blog or answer. Free, no sign-up; the credit line links back here.

Embed
Cite

Found this useful? Share it

Share

Founder · Cloud & Infrastructure

0xLax

I'm Laxman. I started Aliteq in Kathmandu in 2019 as a PC hardware shop — building, repairing and selling machines — and ran it until 2024. These days I live in Denmark and build AI apps, and aliteq.com is where I work out, in public, what things actually cost: servers, clouds, GPUs, and the software a small company ends up paying for. I write the infrastructure pieces myself because I've had the screwdriver in my hand.

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading