Skip to content
HomeHome
DE
WhatsAppMailPhone
← All articles
Six EU-hosted AI APIs compared on price, contracts and speed
KI

Six EU-hosted AI APIs compared on price, contracts and speed

Photo: Brett Sayles / Pexels

IONOS AI Model Hub, STACKIT AI Model Serving, Hetzner Inference, Scaleway, OVHcloud and Mistral compared as of 21 September 2026. Prices per million tokens, data centre location, GDPR data processing agreement, SLA and my own measurement of the same model on Hetzner and Scaleway.

Eric MengeAuthorEric MengeOwner & web developer at EMIT Solution
Published
Reading timeca. 9 min

In short

  • The same model costs very different amounts across EU providers. On 21 September 2026 the output price for Qwen3.8-27B ranged from 0.65 euros at STACKIT to 3.30 euros at Scaleway per million tokens, so the most expensive offer is about five times the cheapest.
  • IONOS AI Model Hub and STACKIT AI Model Serving both process requests in Germany and offer a GDPR data processing agreement plus a committed availability of 99.9 and 99.5 percent respectively.
  • Hetzner Inference is free but runs as an experiment without an SLA and without a stated data processing agreement. That rules it out for customer data in production.
  • In my measurement on 21 September 2026 with Qwen3.6-35B-A3B, Scaleway started answering after 0.49 seconds median, Hetzner after 2.3 seconds. On throughput there was no clear winner at 43 and 50 tokens per second.

If you want to build a language model into your own application via an API and keep the data out of the United States, there is now a real choice of providers. On 21 September 2026 I compared six providers with data centres in Germany or elsewhere in the EU, based solely on their own pricing and documentation pages. I also ran a small measurement, giving Hetzner and Scaleway the same model and the same task.

The question behind it is the one a business actually asks. Where do I get a model that computes in the EU, with a GDPR data processing agreement, at what price and which of these options can carry a production workload?

One detail makes the comparison easier. Qwen3.8-27B, a mid-sized open language model with 27 billion parameters, is in the catalogue of five of the six providers. That makes it possible to compare the price of exactly the same model side by side.

The six providers at a glance

All six offer an OpenAI-compatible API. Requests have the same shape as with OpenAI, so existing code usually only needs a different base URL, a different key and a different model name. Prices are list prices per 1 million tokens for input and output, as of 21 September 2026:

Provider Location Qwen3.8-27B gpt-oss-120b DPA SLA
IONOS AI Model Hub Germany €0.40 / €2.70 €0.15 / €0.65 yes, part of the terms 99.9%
STACKIT AI Model Serving Germany €0.45 / €0.65 €0.45 / €0.65 yes, separate agreement 99.5%
OVHcloud AI Endpoints France €0.40 / €2.70 €0.08 / €0.40 yes, annex to the terms 99.98%
Scaleway Generative APIs France €0.60 / €3.30 €0.15 / €0.60 yes, standard agreement 99.9%
Mistral La Plateforme EU, US optional not offered not offered yes, part of the contract enterprise only
Hetzner Inference EU free not offered not stated none

DPA stands for data processing agreement, the contract required by Article 28 of the GDPR (the EU General Data Protection Regulation) in which the provider commits to processing personal data only on your behalf. The SLA is the contractually committed monthly availability.

When it comes to what happens with the content, the providers are close to each other. IONOS, STACKIT, OVHcloud, Scaleway and Hetzner all state that they do not store prompts and responses permanently and do not use them for training. The differences are in the details.

The same model costs up to five times as much

For Qwen3.8-27B the output price ranges from 0.65 euros at STACKIT to 3.30 euros at Scaleway. Scaled to a realistic volume, say 10,000 requests with 1,000 input tokens and 400 output tokens each, that comes to 7.10 euros at STACKIT, 14.80 euros at IONOS and OVHcloud and 19.20 euros at Scaleway. For gpt-oss-120b the order almost flips. The same volume costs 2.40 euros at OVHcloud, 3.90 euros at Scaleway, 4.10 euros at IONOS and 7.10 euros at STACKIT.

The reason for this pattern is STACKIT’s pricing. It does not price per model but in three tiers. Everything in the Plus tier costs the same, whether it is Qwen3.8-27B, Llama 3.3 70B or gpt-oss-120b. That is cheap for models that are expensive elsewhere and less attractive for models that are cheap elsewhere. No provider wins across the board. It pays to pick the model first and the provider second.

The biggest cost lever, though, does not appear in any price table. Many current models think before they answer, and that thinking is billed as output. At Scaleway this is switched on by default for Qwen3.6-35B-A3B. The setting that turns off thinking at Hetzner was ignored by Scaleway in my tests. Five requests each ran up to the limit of 800 output tokens and came back without a single line of answer. Only this parameter made the model answer directly:

{
  "model": "qwen3.6-35b-a3b",
  "reasoning_effort": "none"
}

After that, the same email needed about 175 instead of 800 output tokens. At list price that is 0.03 instead of 0.12 cents per request. Only the cheaper variant produces a result at all. gpt-oss-120b is a reasoning model too, so its real consumption is higher than the visible answer suggests.

Compass lying on a map of Europe Photo: Aliaksei Lepik / Pexels

The providers one by one

IONOS AI Model Hub

IONOS, a large German hosting company, offers eight language models, a broad range of well-known open models from the small Llama 3.1 8B through Mistral Small 24B to the large Qwen3.5-397B. Inference runs in Germany. According to the documentation, prompts and responses are never written to logs or disk and only sit in host and GPU memory during the request. The data processing agreement has been part of the IONOS terms since July 2022, so no separate signature is needed. IONOS commits to 99.9 percent availability as a monthly average. I did not find a free quota on IONOS’s pages. Prices are listed in the IONOS price list.

STACKIT AI Model Serving

STACKIT, the German cloud operated by Schwarz Digits, runs in its Germany South region and offers six active language models, including Qwen3.8-27B, Llama 3.3 70B, gpt-oss-120b and Gemma 4 31B. According to its FAQ, requests are neither stored nor used for training. The service certificate adds that the user’s email address and subject ID stay in the service logs for 30 days. Availability is committed at 99.5 percent per calendar month. One rule matters for planning. STACKIT may retire a model with six months’ notice, or with three months’ notice when a direct successor is released. There is a 30-day trial, and STACKIT issues access within 24 hours after a manual check.

OVHcloud AI Endpoints

For gpt-oss-120b, OVHcloud is the cheapest provider in this comparison at 0.08 euros input and 0.40 euros output. The selection of pure text models is small, supplemented by several Qwen models that also understand images, Qwen3.8-27B among them. According to the documentation the infrastructure is located in Gravelines, France. OVHcloud states that data is never used for training and that it keeps only what is needed for billing. For the standard API it lists 99.98 percent availability, while the batch API is still in beta without a commitment. You can even test without an account, limited to two requests per minute and model. Prices are in the model catalogue.

Scaleway Generative APIs

On 21 September 2026, the API catalogue of the French provider Scaleway listed sixteen models, including Llama 3.3 70B, Mistral Small 3.2, Mistral Medium 3.5, gpt-oss-120b, DeepSeek-V4-Flash, GLM-5.2 and several Qwen variants. Inference runs in Paris. According to its privacy page Scaleway does not store content, with one exception. If traffic disrupts operations, for example through server errors, the full request may be kept for up to two weeks for troubleshooting. No training takes place. The first 1,000,000 tokens are free, and the SLA states 99.9 percent with service credits if it is missed. For Qwen3.8-27B Scaleway is the most expensive provider, while for gpt-oss-120b and for Mistral Small 3.2 at 0.15 and 0.35 euros it is very cheap.

Mistral La Plateforme

Mistral’s own platform offers almost exclusively Mistral models. A mid-sized model such as Mistral Small 4 costs 0.15 dollars input and 0.60 dollars output, Mistral Large 3 comes in at 0.50 and 1.50 dollars. Data is hosted in the EU by default, and only if you explicitly choose the US endpoint does it end up in the United States. The data processing agreement becomes part of the contract automatically. Two points deserve a look. The admin panel has a toggle that stops API calls from being used to improve Mistral’s services. And zero data retention, the commitment not to keep inputs and outputs at all, is only available on request for pay-as-you-go customers. Mistral provides an SLA for enterprise customers only.

Hetzner Inference

Hetzner is the special case. Its inference API is a free experiment with two models, Qwen3.6-35B-A3B and Qwen3.8-27B. According to the documentation, Hetzner stores neither the request nor the response. Each key is allowed ten requests per minute, plus four million input and 100,000 output tokens. Hetzner does not state a data processing agreement for the experiment and there is no SLA. Hetzner itself advises against building production environments on it. For development this is an excellent offer, for customer data the contractual basis is missing.

Stopwatch app on a smartphone Photo: Castorly Stock / Pexels

Hetzner and Scaleway measured with the same model

Hetzner and Scaleway both offer Qwen3.6-35B-A3B, one for free, the other for 0.25 and 1.50 euros per million tokens. On the afternoon of 21 September 2026 I gave both the same task. A casual message to a customer about a delayed delivery was to be rewritten as a formal German business email. Each provider got two rounds of five requests, about 25 minutes apart. I measured two values. Time to first token shows how long it takes until any text arrives. Throughput shows how fast text is produced after that.

Median of 10 runs Hetzner Scaleway
Time to first token 2.3 s 0.49 s
Range 1.4 to 6.3 s 0.28 to 1.8 s
Throughput 50 tokens/s 43 tokens/s
Range 32 to 74 tokens/s 19 to 62 tokens/s
Total time per email 6.0 s 4.6 s

Scaleway starts writing noticeably sooner. The network is not the reason, since setting up the connection took about 150 milliseconds to both servers. The waiting happens on the server side. On throughput there is no clear winner, as the values fluctuated a lot for both. At Hetzner throughput halved between the two rounds, from 72 to 36 tokens per second, while the wait for the first token rose from 1.7 to 3.1 seconds. That fits a free service shared by many users.

Language quality was fine on both. None of the 20 emails contained stray foreign characters, and all had a subject line, a greeting and a clean closing.

I also tried Qwen3.8-27B on both. At Scaleway the first token arrived after about half a second. At Hetzner three answers only arrived after 141 to 180 seconds, each in a single block, and the fourth request hit my 180-second timeout. At the same time Qwen3.6 on the same endpoint answered normally. This is exactly what Hetzner means when it calls the service an experiment.

The whole measurement used about 10,000 tokens at Scaleway and stayed entirely within the free quota. At list price it would have cost a little over one cent.

Hand signing a document Photo: Tima Miroshnichenko / Pexels

An EU data centre alone does not make processing GDPR compliant

Location is an important building block, but only one. As soon as personal data goes into the requests, such as names, email addresses or customer messages, you need a data processing agreement with the provider, a legal basis for the processing and a matching section in your privacy policy. It is also worth checking the sub-processors. Mistral, for example, notes that data may temporarily go to sub-processors outside the EU that are listed in its Trust Center.

For Hetzner Inference this means only test data without any personal information belongs there as long as no agreement exists. How the large US providers handle stored prompts and what zero data retention means there is covered in a separate article.

Which provider for which purpose

For trying things out, Hetzner Inference is hard to beat as long as no real customer data is involved. If you want to test with the model that will later run in production, use Scaleway’s free quota, STACKIT’s 30-day trial or OVHcloud’s free access.

For production with customer data and a requirement to stay in Germany, IONOS and STACKIT remain. Both provide a data processing agreement, a committed availability and inference in Germany. STACKIT is the cheaper choice for Plus-tier models with a lot of output, while IONOS has the wider selection, especially of small and inexpensive models such as Mistral Small 24B at 0.10 and 0.30 euros.

For cheap high-volume work, such as pre-sorting thousands of documents, OVHcloud is hard to ignore for gpt-oss-120b. Scaleway follows closely with gpt-oss-120b and Mistral Small 3.2. In both cases the thinking mode should be set deliberately, otherwise it eats up the savings.

If you are weighing which of these providers fits your application and how to integrate the model cleanly into your workflows, feel free to get in touch.

FAQ

Which AI APIs run in a data centre in Germany?+

As of 21 September 2026, IONOS AI Model Hub and STACKIT AI Model Serving state that they process requests in Germany. Hetzner Inference runs in Hetzner's own data centres in the EU without naming a specific site. OVHcloud computes in Gravelines according to its documentation and Scaleway in Paris, both in France. Mistral hosts in the EU by default and also offers a US endpoint.

Is there a free AI API with servers in the EU?+

Yes, several. Hetzner Inference is completely free while the experiment runs, Scaleway does not charge for the first 1,000,000 tokens, OVHcloud allows two requests per minute and model without a key and STACKIT offers a 30-day trial. Mistral lists 10 dollars of API credit per month in its free plan. That is enough for testing, but customer data calls for a paid plan with a contract.

Does an EU data centre make an AI API GDPR compliant?+

Not on its own. As soon as personal data goes into the prompts, you need a data processing agreement with the provider under Article 28 GDPR, a legal basis for the processing and matching information in your privacy policy. The location helps, because it avoids transfers to third countries, but it is one building block among several.

What does a single AI request cost at an EU provider?+

Usually a fraction of a cent. Rewriting a business email with about 150 input tokens and 175 output tokens cost roughly 0.03 cents at Scaleway list price with Qwen3.6 in my measurement. With the model's thinking mode left on, the same request used 800 output tokens and still returned no answer. Output often costs several times as much as input, so that setting matters.

Want to know more?

In a free intro call we discuss how you can use these topics for your company. Not a sales pitch, but an honest assessment.

Book a free intro call