Should a 10–50 person company buy a GPU server, rent one by the hour, or just pay per token? The three-year bill for each, from live rental prices, today's hardware prices and government power statistics. The answer for most small companies is not the one GPU vendors sell.
I spent five years building, repairing and selling PCs at Aliteq's shop in Kathmandu, from 2019 to 2024. These days I build AI apps from Denmark. So when a small company asks whether it should buy its own AI server, I understand why the idea is attractive: you pay once, the machine is yours, and your data never leaves the office.
The numbers usually say otherwise, and the pages that rank for this question won't tell you, because every one of them sells one side of the answer. Lenovo's TCO whitepaper sells servers. Runpod's and Spheron's sell rented GPUs. Lenovo's smallest configuration starts at two RTX PRO 6000 cards for about $68,000, and none of them model a company with one box, a handful of users, or an API bill. aliteq sells neither hardware nor cloud. We do have referral links with Vast and Runpod, disclosed wherever they appear, and they don't change the maths here.
What we do have is our own data. aliteq has logged Runpod's and Vast's GPU prices every hour since 21 July 2026, so the rental numbers below are ours, not a vendor's. The hardware prices were read on retailer and builder pages on 27 September 2026, and the power prices come from government statistics.
The answer in one chart
For a typical small company, the API is the cheapest option by a wide margin, renting a GPU by the hour comes next, and owning comes last. Over three years, a 25-person internal document chat costs $356 to $5,148 through an API, about $10,391 on a rented RTX PRO 6000 in business hours, and about $32,029 on an owned one.
The same workload, priced six ways over three years. Assumptions are listed below the chart and in the calculator. · aliteq research
How much AI work does a small company actually do?
Far less than GPU vendors assume. Their break-even maths assumes a box busy 60–70% of the time. By my estimate, a 25-person document chat needs about four hours of real GPU work a month on one RTX PRO 6000, and a month of batch document processing about two. A coding assistant for ten developers is the exception.
No vendor publishes how many tokens a chat question or a document page uses, so these three workloads are my assumptions, stated in full:
Three workloads (my assumptions)
Document chat, 25 staff
Assumption
20 questions a workday each; 4,000 tokens in, 500 out per question
Tokens per month
44M in + 5.5M out
GPU time a month (1× RTX PRO 6000)
~4 hours of work
Coding assistant, 10 developers
Assumption
20M input tokens per developer-day, 90% of them cached; 150k out
Tokens per month
4,400M in (3,960M cached) + 33M out
GPU time a month (1× RTX PRO 6000)
~30–140 hours
Document processing
Assumption
20,000 pages; 800 tokens in and 200 out per page
Tokens per month
16M in + 4M out
GPU time a month (1× RTX PRO 6000)
~2 hours
Assumption
Tokens per month
GPU time a month (1× RTX PRO 6000)
Document chat, 25 staff
20 questions a workday each; 4,000 tokens in, 500 out per question
44M in + 5.5M out
~4 hours of work
Coding assistant, 10 developers
20M input tokens per developer-day, 90% of them cached; 150k out
4,400M in (3,960M cached) + 33M out
~30–140 hours
Document processing
20,000 pages; 800 tokens in and 200 out per page
16M in + 4M out
~2 hours
The coding profile is calibrated to the one real figure I could find: Anthropic's documentation reports that Claude Code costs enterprises about $13 per developer per active day on average. My assumption works out to about $9 a day on Sonnet 5 and $15 on Opus 5.5, which brackets it.
The GPU-time column uses published benchmarks, not tests of our own: figures in llama.cpp's gpt-oss guide show about 4,500 tokens a second of prompt processing for gpt-oss-120b on one RTX PRO 6000. The point isn't the exact hours. It's that a small company's chat and document work keeps a dedicated GPU idle nearly all day.
What the API and seats cost over three years
Per-token pricing makes the chat and document workloads almost free on open models and a few thousand dollars on frontier ones. Coding is where the bill grows: tens of thousands of dollars on API pricing, or a few thousand on per-seat coding plans.
3-year cost by model (list prices, 27 Sep 2026)
gpt-oss-120b, hosted ($0.15 in / $0.60 out per 1M)
Document chat, 25 staff
$356
Coding, 10 developers
$24,473
Documents (batch)
$173
gpt-6-luna ($0.10 / $0.50)
Document chat, 25 staff
$257
Coding, 10 developers
$3,604
Documents (batch)
$65
Claude Haiku 4.5 ($1 / $5)
Document chat, 25 staff
$2,574
Coding, 10 developers
$36,036
Documents (batch)
$648
Claude Sonnet 5 or gpt-6-sol ($2 / $10)
Document chat, 25 staff
$5,148
Coding, 10 developers
$72,072
Documents (batch)
$1,296
Claude Opus 5.5 ($4 / $20)
Document chat, 25 staff
$10,296
Coding, 10 developers
$115,632
Documents (batch)
$2,592
Document chat, 25 staff
Coding, 10 developers
Documents (batch)
gpt-oss-120b, hosted ($0.15 in / $0.60 out per 1M)
$356
$24,473
$173
gpt-6-luna ($0.10 / $0.50)
$257
$3,604
$65
Claude Haiku 4.5 ($1 / $5)
$2,574
$36,036
$648
Claude Sonnet 5 or gpt-6-sol ($2 / $10)
$5,148
$72,072
$1,296
Claude Opus 5.5 ($4 / $20)
$10,296
$115,632
$2,592
Seats change the picture for people rather than systems. Twenty-five ChatGPT Business or Claude Team seats cost $18,000 over three years at $20 a month billed annually, but they include the chat app, admin controls and SSO that a raw API doesn't. For ten developers, GitHub Copilot Business comes to $6,840 over three years and Claude Team standard seats, which include Claude Code, to $7,200.
What owning actually costs
A box's three-year bill is its price, plus power, plus someone's time, minus what you sell it for. For one RTX PRO 6000 workstation in business hours that's about $32,029: $26,979 for the machine, $654 of US power, $9,792 of admin time at four hours a month, less a 20% resale credit.
Own vs rent, 3 years (US power, business hours, 4 h/month admin, 20% resale)
DIY box, 1× RTX 5090 ($7,300)
Own, 3 years
$16,435
Rent the same card, business hours (Runpod)
$4,372
Rent 24/7 (Runpod)
$18,133
1× RTX PRO 6000 Max-Q ($26,979)
Own, 3 years
$32,029
Rent the same card, business hours (Runpod)
$10,391
Rent 24/7 (Runpod)
$43,099
2× RTX PRO 6000 Max-Q ($51,858)
Own, 3 years
$52,333
Rent the same card, business hours (Runpod)
$20,782
Rent 24/7 (Runpod)
$86,198
4× RTX PRO 6000 Max-Q (~$84,600)
Own, 3 years
$79,192
Rent the same card, business hours (Runpod)
$41,564
Rent 24/7 (Runpod)
$172,397
NVIDIA DGX Spark ($4,699)
Own, 3 years
$13,776
Rent the same card, business hours (Runpod)
not tracked
Rent 24/7 (Runpod)
not tracked
Strix Halo mini-PC ($3,500)
Own, 3 years
$12,714
Rent the same card, business hours (Runpod)
not tracked
Rent 24/7 (Runpod)
not tracked
Own, 3 years
Rent the same card, business hours (Runpod)
Rent 24/7 (Runpod)
DIY box, 1× RTX 5090 ($7,300)
$16,435
$4,372
$18,133
1× RTX PRO 6000 Max-Q ($26,979)
$32,029
$10,391
$43,099
2× RTX PRO 6000 Max-Q ($51,858)
$52,333
$20,782
$86,198
4× RTX PRO 6000 Max-Q (~$84,600)
$79,192
$41,564
$172,397
NVIDIA DGX Spark ($4,699)
$13,776
not tracked
not tracked
Strix Halo mini-PC ($3,500)
$12,714
not tracked
not tracked
Change any of it here. The defaults are the numbers above, and the admin time, resale value and hours are sliders because they're my assumptions:
$1.64/h × 1 GPU for the same hours · Vast median: $9,753
Break-even
18.1 h/day
Hours a day, every day for 3 years, before owning beats renting the same card.
For comparison: 25 staff using an internal document chat through an API costs about $356 (hosted open model) to $5,148 (Claude Sonnet 5 / gpt-6-sol) over the same three years, on our stated token assumptions.
Hardware: Bizon, Puget, NVIDIA and GMKtec pages; rental: aliteq GPU tracker (Runpod on-demand, Vast 30-day median); power: EIA, Eurostat (excl. VAT and recoverable taxes), DESNZ; admin wage: BLS. All checked 27 Sep 2026. Admin time, resale, usage and the 4-GPU price scaling are our assumptions. Power excludes cooling; colocation not included. An estimate, not a quote.
Three things about that table surprise people:
Admin time is the biggest hidden line. Nobody measures how long one small inference box takes to look after, so I use four hours a month, below the 10–20 hours some guides assume for production serving. At a loaded US sysadmin cost of about $68 an hour (the BLS median wage of $47.66, divided by the 70% wages make up of total pay), that's $9,792 over three years. For a DGX Spark or a Strix Halo box, admin is 71–77% of the whole bill.
Power barely matters. It's 2–17% of the cost of owning in every case I modelled. One RTX PRO 6000 box costs about $654 in US power over three years in business hours, $1,027 in Denmark and $2,190 in the UK, using government statistics for business electricity.
Hardware is expensive right now. The RTX 5090's $1,999 launch price can't be paid anywhere I checked: partner cards list at $4,400–$4,900 and are sold out, and an in-stock one was $7,980 at Best Buy. An RTX PRO 6000 card is $11,739–$16,269 and backordered. The DGX Spark is $4,699 and the GMKtec Strix Halo box with 128 GB is $3,499.99.
On taxes, owning has one real advantage: in the US, IRS Publication 946 treats computers as five-year property, and for tax years beginning in 2026 the Section 179 deduction goes up to $2,560,000, so a company can often write off a server in the year it buys it. In Denmark, equipment depreciates at up to 25% a year on a declining balance. Ask your accountant; this isn't tax advice.
How many hours a day your server has to work
To beat renting the same card on Runpod, one RTX PRO 6000 box has to work about 18 hours a day, every day, for three years. A DIY RTX 5090 box needs 23.5 hours, which is never in practice. Only a four-GPU box gets close to a normal working day, at about 11 hours.
Left end: admin time is free. Right end: four hours a month at $68 an hour. · aliteq research
If someone looks after the box for free, because it's their hobby or already their job, the numbers improve a lot: 8.8 hours a day for the RTX 5090 box and 12.5 for one RTX PRO 6000. That's why admin time, not the GPU price or the power bill, decides the single-box case. Be honest with yourself about who will patch it, update the drivers and answer "the AI is down" at 9 in the morning.
Where to rent, if you rent
Rent from a specialist GPU cloud, not a hyperscaler. One H100 costs $6.88 an hour on AWS against $2.69 on Runpod and a $2.18 median on Vast. For a small company's private open model, a 96 GB RTX PRO 6000 at about $1.64 an hour is usually enough.
Snapshot from aliteq's GPU price tracker, 27 September 2026; AWS on-demand price feed. · aliteq research
Our tracker also contradicts a story you may have seen, that GPU rental prices have doubled. Since we started logging on 21 July 2026, Runpod's list prices haven't changed at all, and Vast's median H100 price has fallen 24%. The one card that rose is the RTX PRO 6000, up 49% on Vast, which is consistent with its purchase price climbing. The live table updates every hour.
If you want a whole machine by the month in the EU instead, Hetzner's GEX131 with one RTX PRO 6000 Max-Q costs €1,197.30 a month plus €599 setup, which is about $49,833 over three years at today's exchange rate. That's more than owning the same card and more than renting it around the clock on Runpod.
When owning is the right call
Own the server when your data must not leave the building, or when a steady workload keeps the box busy most of every day. A coding assistant for a large team, round-the-clock batch jobs, or a four-GPU box that's used all day can justify it. A chat assistant for 25 people can't.
The data reason is real and often decisive. If a customer contract, a regulator or your own risk appetite rules out sending documents to a third party, a local model on your own hardware removes the question entirely. The Security & Compliance section covers what customers actually ask for, and SOC 2 for AI startups covers what your API vendors already cover.
If you do buy, plan the room before the box. Every watt a server draws becomes heat: about 3,412 BTU an hour per kilowatt, according to the EIA. Four 600-watt cards need more power than a normal office circuit is built for, so check with an electrician first. Colocation is the alternative, but cheap single-server deals include only 60–360 watts. FDC Servers' high-density offer, $999 a month with 4 kW included, is the kind of plan a four-GPU box needs.
What I'd do at a 10–50 person company
Start with an API or a handful of seats. If you need a private open model, rent the GPU by the hour until you can see a steady, all-day workload in your own usage logs. Buy hardware only when data control demands it or your logs prove the box would be busy.
Is it cheaper to buy an AI server or rent cloud GPUs for a small company?
Usually renting, and often an API is cheaper still. One RTX PRO 6000 box costs about $32,029 over three years including admin and resale, against about $10,391 to rent the same card on Runpod for business hours. Owning only wins if the box works about 18 hours a day for three years.
How many users can one GPU server handle?
There's no fixed number. It depends on the model, prompt length and how many people ask at the same moment. For a small company's document chat, one RTX PRO 6000 has far more capacity than 25 people use; the real limit is usually how many ask at once, not the total.
How much does it cost to run an AI server in electricity?
Less than people expect. One RTX PRO 6000 workstation costs about $654 over three years in US business electricity during business hours, about $1,565 if it runs flat out around the clock. Power is 2–17% of the total cost of owning in every case we modelled.
Is self-hosting an LLM cheaper than using the OpenAI or Claude API?
For small companies, rarely. A 25-person document chat costs about $356 over three years on a hosted open model and $5,148 on Claude Sonnet 5 or gpt-6-sol, while owning a capable box costs over $12,000 once admin time is included. Self-hosting pays off for data control or very heavy, steady use.
What is a GPU server worth after three years?
In a normal market, a fraction of its price; I assume 20%. Today's shortage has pushed used prices unusually high: a used RTX 4090 sells for more than its launch price. Don't count on that lasting.
Should a small company use AWS for GPUs?
Not for simple GPU rental. AWS charges $6.88 an hour for one H100, against $2.69 on Runpod. Hyperscalers make sense when you need their other services, compliance paperwork or existing contracts.
The rest of the Cloud & IT Infrastructure section prices the other decisions a small company faces: managed IT, backups, device management and the Windows Server 2016 deadline.
Use this in your own page
Teaching this? Paste the live version into your course, blog or answer. Free, no sign-up; the credit line links back here.