aliteq.

Private LLM cost for EU companies (2026): EU API tiers vs open models vs your own GPU

Your company wants AI on internal documents without the prompts leaving the EU. We priced every route on the vendors' own pages: OpenAI's and Azure's EU tiers, Claude through AWS and Google, Mistral, French and German open-model APIs, and a GPU of your own. The EU premium is 0% to 75%, and for most companies the cheapest private option isn't a server.

TensorUpdated 1h ago14 min readWeb story
Hand-drawn editorial illustration of a desk lamp over a sealed, padlocked glass jar holding a glowing chat bubble, inside a small lime-green fence
Share

"Can we use AI on our own documents without the data leaving the EU?" is the question that stalls a lot of company AI projects. The answer is yes, several ways. What nobody lays out is what each way costs, because the prices sit on a dozen vendors' pages in different units, and some of the most important ones are only in machine-readable price feeds.

So we read them all on 27 September 2026. For AWS and Microsoft, whose price pages load in JavaScript, we pulled the numbers from their own official price APIs. The GPU rental rates come from aliteq's own tracker, which has logged Runpod's and Vast's prices every hour since 21 July 2026. We didn't deploy or benchmark any of these routes, and none of these vendors pay us.

What a private LLM costs a month, route by route

At 10 million tokens a workday, the cheapest way to keep prompts in the EU is an open model on a French cloud: gpt-oss-120b on OVHcloud costs about $29 a month. The same volume on a frontier model's EU tier runs $350 for Claude Haiku 4.5 up to $763 for GPT-6 on Azure's EU Data Zone.

Bar chart of the monthly cost of 10 million tokens a workday for an EU company, with data kept in the EU: gpt-oss-120b on OVHcloud $29, gpt-6-luna OpenAI EU $35, Scaleway $50, Bedrock Frankfurt $58, Mistral Large 3 $134, Claude Haiku 4.5 $350, rented Runpod GPU $360, Gemini 3.8 Flash EU $524, rented OVHcloud H100 $634, gpt-6-sol OpenAI EU $699, Claude Sonnet 5 EU $699, gpt-6-sol Azure EU Data Zone $763, own server $914.
Thirteen EU-resident routes, one workload. Violet: open models. Coral: frontier models' EU tiers. Grey: your own GPU. · aliteq research

How big is 10 million tokens a day? A 25-person team asking an internal assistant 20 questions a day each, with about 4,000 tokens of documents and question going in and 500 tokens coming back, uses about 2.25 million. So 1 million a day is a small team's light use, 10 million is heavy company-wide use, and 100 million is a product serving customers.

Monthly bill at three volumes (22 workdays, 8:1 input to output, no caching)

gpt-6-sol, global (not EU-resident, for reference)

1M tokens a day
$64
10M a day
$636
100M a day
$6,356

gpt-6-sol, OpenAI EU residency

1M tokens a day
$70
10M a day
$699
100M a day
$6,991

gpt-6-sol, Azure EU Data Zone

1M tokens a day
$76
10M a day
$763
100M a day
$7,627

Claude Sonnet 5, Bedrock Frankfurt or Vertex EU

1M tokens a day
$70
10M a day
$699
100M a day
$6,991

Gemini 3.8 Flash, EU (2027 price)

1M tokens a day
$52
10M a day
$524
100M a day
$5,243

Mistral Medium 3.5, EU-hosted

1M tokens a day
$48
10M a day
$477
100M a day
$4,767

Claude Haiku 4.5, Bedrock Frankfurt

1M tokens a day
$35
10M a day
$350
100M a day
$3,496

Mistral Large 3, EU-hosted

1M tokens a day
$13
10M a day
$134
100M a day
$1,344

gpt-oss-120b, Bedrock Frankfurt

1M tokens a day
$5.84
10M a day
$58
100M a day
$584

gpt-oss-120b, Scaleway (Paris)

1M tokens a day
$5.02
10M a day
$50
100M a day
$502

gpt-6-luna, OpenAI EU residency

1M tokens a day
$3.50
10M a day
$35
100M a day
$350

gpt-oss-120b, OVHcloud (France)

1M tokens a day
$2.90
10M a day
$29
100M a day
$290

Rent one GPU (Runpod EU or OVHcloud H100)

1M tokens a day
$360–634
10M a day
$360–634
100M a day
$1,269–2,284 (24/7)

Own server, 1× RTX PRO 6000, 3 years

1M tokens a day
$914
10M a day
$914
100M a day
$914 + extra power

The formula is monthly tokens times the blended price: eight parts input at the input price plus one part output at the output price, divided by nine. Two caveats change the table in practice. Claude's newer models count about 30% more tokens for the same text than other models do, so Claude's real bill per word is higher than a per-token table suggests. And Gemini 3.8 Flash has an introductory price until the end of 2026 that is half the 2027 price shown here; don't budget a year on the intro price.

Try your own volume:

What a private LLM costs a month

Cheapest EU route

$28.94/mo

gpt-oss-120b · OVHcloud (France)

Cheapest frontier model, EU

$34.96/mo

gpt-6-luna · OpenAI EU

Cheapest own GPU

$360/mo

Rent RTX PRO 6000 · Runpod EU

  • gpt-oss-120b · OVHcloud (France)Open model, EU API$28.94
  • gpt-6-luna · OpenAI EUFrontier model, EU tierEU residency needs sales approval$34.96
  • gpt-oss-120b · Azure EU Data ZoneOpen model, EU API$48.40
  • gpt-oss-120b · Scaleway (Paris)Open model, EU API$50.16
  • gpt-oss-120b · Bedrock FrankfurtOpen model, EU API$58.42
  • Mistral Large 3 · MistralOpen model, EU APIswitch the training opt-out off yourself$134
  • Claude Haiku 4.5 · Bedrock FrankfurtFrontier model, EU tier$350
  • Rent RTX PRO 6000 · Runpod EUYour own GPUbusiness hours · EU data-centre price not verified; tracker rate used$360
  • Mistral Medium 3.5 · MistralFrontier model, EU tier$477
  • Gemini 3.8 Flash · Google EU (2027 price)Frontier model, EU tier$524
  • Rent H100 · OVHcloudYour own GPUbusiness hours$634
  • gpt-6-sol · global (not EU-resident)Not EU-residentreference only: data may leave the EU$636
  • gpt-6-sol · OpenAI EUFrontier model, EU tierno Batch discount with EU residency$699
  • Claude Sonnet 5 · Bedrock or Vertex EUFrontier model, EU tier$699
  • Managed H100 · ScalewayYour own GPUbusiness hours · model served for you$754
  • gpt-6-sol · Azure EU Data ZoneFrontier model, EU tier$763
  • Own RTX PRO 6000 server · 3 yearsYour own GPUbusiness hours · more power at 24/7 not included$914
  • Hetzner GEX131 · monthlyYour own GPUbusiness hours · + $683 setup once$1,437

List prices read 27 Sep 2026 on each vendor's page or official price feed (OpenAI, Azure, AWS, Google Cloud, Mistral, Scaleway, OVHcloud), ex-VAT, EUR converted at the ECB rate of 25 Sep 2026. No caching or batch discounts. Own-GPU rows run gpt-oss-120b on one 80–96 GB GPU, which serves about 99M tokens in an 8-hour day in theory; real traffic peaks leave less. Rental rates from aliteq's GPU price tracker. Admin time and utilization are our assumptions. An estimate, not a quote.

The EU premium isn't one number

Keeping inference in the EU costs anywhere from nothing to 75% more than the same model's global price. For the big frontier APIs it's usually 10%, and for GPT-6 on Azure it's 20%.

Bar chart of the extra cost of keeping AI inference in the EU: Mistral API 0%, OpenAI EU residency +10%, Claude on Bedrock or Vertex EU +10%, Gemini 3.8 Flash non-global +10%, Azure EU Data Zone gpt-oss-120b +10%, Azure EU Data Zone GPT-6 +20%, gpt-oss-120b on Bedrock Frankfurt +33% vs US, Mistral Enterprise APIs +75%.
What the EU tier adds, vendor by vendor, from each vendor's own page or price feed. · aliteq research

What each vendor's own pages say:

  • OpenAI charges a 10% uplift for regional processing on models released since 5 March 2026. EU residency isn't self-serve: you contact sales and must be approved for abuse-monitoring controls. For gpt-6-sol and gpt-6-luna it's available only with Standard processing, so the 50% Batch discount doesn't apply.
  • Azure's EU Data Zone adds 20% on GPT-6 but 10% on GPT-5.6 and on gpt-oss-120b, according to Microsoft's own price feed. It's self-serve inside Azure.
  • Claude's own API has no EU option. Its inference can run globally or in the US only. To keep Claude in the EU you go through AWS Bedrock in an EU region or Google Cloud's EU multi-region, both at a 10% premium. Sonnet 5 costs $2.20 in and $11 out per million tokens there, against $2 and $10 globally.
  • Mistral hosts data in the EU by default at its list price. Its Enterprise APIs, which add regional data processing controls and SLAs, cost 75% above list.
  • Open models are the surprise. gpt-oss-120b on OVHcloud in France is 34% cheaper than the common US price for the same model. On Bedrock in Frankfurt it's 33% dearer than on Bedrock in the US. There's no single EU premium for open weights.

One more catch: Bedrock's region table tags Zurich and London as "EU" regions, and neither is in the EU. If you need EU-only, pin a region in an EU member state.

When your own GPU pays off

Running your own GPU only beats an API when the alternative is a frontier model. Against gpt-6-sol's EU tier, one rented GPU running gpt-oss-120b in business hours breaks even at about 5 million tokens a workday, and an owned server at about 13 million. Against an EU open-model API, it never breaks even at a volume one GPU can serve.

Bar charts of the tokens per workday at which self-hosting gpt-oss-120b breaks even: against gpt-6-sol with OpenAI EU residency, rented Runpod GPU 5.2 million, rented OVHcloud H100 9.1 million, own server 13.1 million; against Claude Haiku 4.5 on Bedrock Frankfurt, 10.3, 18.1 and 26.1 million.
Break-even is tokens a workday at which the monthly self-host cost equals the API bill. · aliteq research

The reason is capacity. One 80–96 GB GPU such as an H100 or an RTX PRO 6000 fits gpt-oss-120b, and from published throughput benchmarks it can serve about 99 million tokens in an 8-hour day in theory. Against OVHcloud's price for the same model, a rented GPU would need 124 million tokens a day to break even, beyond what it can serve. So the real choice isn't "API or server". It's which model you need:

  • If an open model is good enough, an EU open-model API is the cheapest private route at every volume a small company has.
  • If you need GPT-6 or Claude quality, you pay the EU tier, and self-hosting only helps if you're willing to swap to an open model at high volume.
  • If the data must never leave your building, you own the hardware, and you're buying control, not savings. Our three-year own-vs-rent calculator prices that choice in full.

The own-server figure here is $914 a month: one RTX PRO 6000 workstation at $26,979 over three years, with four hours of admin a month and 20% resale, using Danish business power and wage rates as the EU example. Which models fit which card is on our cost to run pages, and live rental prices are on our GPU price table. Our tracker stores one price per GPU type, not per region, so check the EU data-centre price when you deploy.

Who keeps your prompts, and where

OpenAI, Azure, Google Cloud and Scaleway say in their docs that they don't train on your API data by default, and OVHcloud says data isn't stored or shared. Mistral's pay-as-you-go wording doesn't say which way its training setting starts, so switch it off explicitly. Zero data retention usually needs an application or a sales conversation.

Table of EU AI routes: OpenAI EU, Azure EU Data Zone, Claude via Bedrock EU, Claude or Gemini on Vertex EU, Mistral, Scaleway and OVHcloud, showing where inference runs, whether each trains on API data by default, zero-retention options, and the EU price versus global.
From each vendor's own data-handling docs, 27 September 2026. · aliteq research

Three details from the docs matter more than the marketing:

  • Residency isn't total. OpenAI's EU residency excludes "system data" such as account, billing and usage metadata, which may be processed outside the region. Mistral says data can be temporarily transferred outside the EU for some features, to listed sub-processors.
  • Bedrock keeps Claude's maker out of the loop. AWS's docs say model providers have no access to Bedrock logs or to customers' prompts and completions.
  • Vendor slogans aren't facts. Scaleway says it isn't subject to the US CLOUD Act, IONOS calls itself 100% GDPR-compliant, and Runpod says it is fully GDPR-compliant in its EU regions. Those are the vendors' own claims. We haven't verified them.

What GDPR actually requires

GDPR doesn't require EU hosting. It requires two things when an AI vendor processes personal data for you: a written processor contract, and a lawful basis for any transfer outside the EU.

  • A processor contract. Article 28 says you may use only processors that give "sufficient guarantees", under a written contract that covers what's processed, for how long and on whose instructions, including transfers. That's the vendor's Data Processing Agreement, and you need it whether the servers are in Paris or Oregon.
  • A transfer basis. Sending personal data outside the EU needs a basis under Chapter V. For US companies that participate in the EU–US Data Privacy Framework, the Commission's adequacy decision of 10 July 2023 lets data flow freely. Check each vendor on the official Data Privacy Framework list; we couldn't read it on 27 September because the list only loads in a browser.
  • Why companies pay for EU residency anyway. The Data Privacy Framework has a live legal challenge: an appeal, case C-703/25 P, was lodged on 31 October 2025, and we found no ruling as of September 2026. EU residency removes the transfer question if the framework falls. It doesn't remove the processor contract.

The EU AI Act adds duties by role. If you put a chatbot in front of customers under your own name, you may be its provider and must tell people they're talking to an AI, a duty that applies since 2 August 2026. An internal document assistant used by staff is usually not high-risk; a system that screens CVs may be, and those rules start on 2 December 2027. The timeline is on our security and compliance hub.

This is a summary of the law and official guidance, not legal advice.

How to choose

Check what you'll send. No personal or confidential client data? Any API will do; pick on price and quality.

If personal data is involved, get the vendor's Data Processing Agreement and check its transfer basis, or choose an EU-resident route.

Test whether an open model is good enough for your task. If it is, an EU open-model API such as OVHcloud or Scaleway is the cheapest private route.

If you need GPT-6 or Claude, budget 10–20% over the global price and ask for EU residency early: OpenAI's needs sales approval.

Switch off any training setting explicitly and ask about zero data retention, which usually needs an application.

Consider your own GPU only for steady volume above about 5–13 million tokens a workday against a frontier model, or when data must stay on your premises.

The full model-by-model table of EU and global prices, and how to switch residency on with each vendor, is in EU data residency for AI APIs.

Quick answers

Does GDPR require us to host AI in the EU?
No. GDPR requires a written processor contract under Article 28 and a lawful basis for any transfer outside the EU, such as the EU–US Data Privacy Framework for participating US companies. Many companies choose EU residency anyway to avoid the transfer question.
How much more does EU data residency cost?
Usually 10% on frontier models: OpenAI's EU residency, Claude on Bedrock or Vertex EU and Gemini's non-global endpoints. GPT-6 on Azure's EU Data Zone is 20% more. Mistral hosts in the EU at list price; its Enterprise APIs with regional controls cost 75% more.
Can I use Claude with EU-only data residency?
Not through Anthropic's own API, which offers global or US inference only. Use Claude through AWS Bedrock in an EU region or Google Cloud's EU multi-region, both at a 10% premium.
What's the cheapest private LLM for an EU company?
An open model through an EU API. gpt-oss-120b on OVHcloud costs about $2.90 a month at 1 million tokens a workday and $29 at 10 million, ex-VAT, at list prices on 27 September 2026.
Is self-hosting an LLM cheaper than an API?
Only against a frontier model, from about 5–13 million tokens a workday. Against an EU open-model API running the same model, one GPU never breaks even at a volume it can serve.
Do these AI APIs train on our data?
By their own docs, OpenAI, Azure, Google Cloud and Scaleway don't train on API data by default, and AWS says model makers can't see Bedrock prompts. Mistral's pay-as-you-go wording is unclear, so switch its training setting off yourself.

Use this in your own page

Teaching this? Paste the live version into your course, blog or answer. Free, no sign-up; the credit line links back here.

Embed
Cite

Found this useful? Share it

Share
Tensor

Local AI & Automation Editor

Tensor

I'm US-based, I run more models at home than I'll admit to, and I've quantized more than I've finished reading about. I write about running AI on your own hardware and, lately, about what it costs a company to do the same — tokens per day, GPUs per month, and the GDPR questions nobody's sales deck answers.

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading