Skip to content
HomeHome
DE
WhatsAppMailPhone
← All articles
Hetzner's free Inference API in n8n needs two changed settings
KI

Hetzner's free Inference API in n8n needs two changed settings

Photo: Christina Morillo / Pexels

Hetzner's free AI API can be connected to n8n through the OpenAI node. In a test with n8n 2.39.9 the default setup failed on the Responses API, and without one extra parameter Qwen used 120 times the tokens it needed. Here is the setup that works.

Eric MengeAuthorEric MengeOwner & web developer at EMIT Solution
Published
Reading timeca. 8 min

In short

  • Hetzner Inference can be connected to n8n with credentials of type OpenAI and the base URL https://inference.hetzner.com/api/v1. In the OpenAI Chat Model node the Use Responses API toggle has to be switched off, otherwise Hetzner answers with a 404 error because it does not offer that interface.
  • Without the parameter chat_template_kwargs with enable_thinking set to false, Qwen reasons before every answer. In a test with n8n 2.39.9 on 21 September 2026, a single email classification cost 480 instead of 4 output tokens and 10.4 instead of 1.2 seconds as a result.
  • Adding /no_think to the prompt does not switch off the reasoning mode of Qwen3.6 or Qwen3.8 on Hetzner. Only the enable_thinking parameter works, and n8n sends it through the Extra Body field of the OpenAI Chat Model node.
  • Of 40 simultaneous requests from n8n to Hetzner Inference, 17 failed with HTTP 429 on 21 September 2026. Retry On Fail did not help. With the batching option of the HTTP Request node, 12 requests spaced 6.5 seconds apart went through without a single error.

Hetzner’s AI interface is still free of charge on 21 September 2026 and it speaks OpenAI’s format. n8n ships an OpenAI node in which the provider’s address can be swapped. That sounds like a five-minute setup. With the default settings, though, I first got an error and then an answer for which the model used 120 times as many tokens as necessary.

So I tested the connection in a fresh n8n installation, version 2.39.9, the current stable release on the day of testing. I checked the OpenAI Chat Model node attached to the Basic LLM Chain and to the AI Agent, the HTTP Request node as an alternative, and how everything behaves at the rate limit. What Hetzner Inference is in general I introduced in July. This article is only about making it run reliably in n8n.

What you need first

You need a key from the experiments area of your Hetzner account and a running n8n instance. The terms are in Hetzner’s documentation. The base URL is https://inference.hetzner.com/api/v1, and two models are on offer, Qwen/Qwen3.6-35B-A3B-FP8 and Qwen3.8-27B. As long as the service is an experiment, it costs nothing. According to Hetzner, the content of requests and responses is not stored, only timestamps and token counts. Per key, a 60-second window allows at most 10 requests, 4 million input tokens and 100,000 output tokens.

Two things are missing, and they decide what you can use this for. There is no commitment on performance or availability. And there is no data processing agreement, the contract that the GDPR (Art. 28) requires before a service provider may process personal data on your behalf. More on that below.

Creating the credentials in n8n

In n8n, create new credentials of type OpenAI. The Hetzner key goes into API Key, and https://inference.hetzner.com/api/v1 goes into Base URL. Leave the Organization ID empty. To test the connection, n8n fetches the model list at that address. Hetzner provides it.

The benefit of this approach shows later. Everything that concerns the provider lives in these credentials. If you move to another provider at some point, you change address and key there, and in the workflow itself only the model name.

Hands on a laptop showing a form with input fields Photo: RDNE Stock project / Pexels

Setting up the model node correctly

For most tasks a Basic LLM Chain is enough, with the OpenAI Chat Model node attached as its model. Under Model choose the ID mode and enter the model name exactly as Hetzner lists it, for example Qwen/Qwen3.6-35B-A3B-FP8.

Right below sits the first trap. The Use Responses API toggle is on by default in n8n 2.39.9. n8n then talks to OpenAI’s newer Responses interface, as the node documentation describes. Hetzner, however, only offers the model list, Completions and Chat Completions. In my test the workflow stopped with the message “The resource you are requesting could not be found”, plus a link to a help page about models that cannot be found. So the message points in the wrong direction, the model name was correct. Toggle off and the request goes through.

Stopping Qwen from overthinking

Both Qwen models run as reasoning models on Hetzner. Before the actual answer they write out a chain of thought that n8n does not show but that costs time and output tokens. I described this for the raw API in my German language test in July. In n8n it plays out like this, measured on a simple task. The model had to assign a made-up complaint email (written in German) to one of five categories:

Setting in the OpenAI Chat Model node Result Output tokens Time in model node
all defaults, Use Responses API on stops with 404 error 0 0.3 s
Use Responses API off “Reklamation” (complaint) 480 10.4 s
plus Maximum Number of Tokens 200 empty text, no error 200 5.1 s
“/no_think” at the end of the prompt “Reklamation” 425 17.9 s
Extra Body with enable_thinking false “Reklamation” 4 1.2 s

The third row is the dangerous one. If you set a token limit to keep answers short, the reasoning eats the whole budget. n8n still reports the run as successful and passes an empty text downstream. In a workflow that branches afterwards, you only notice when items end up in the wrong branch.

Adding /no_think to the prompt, which Qwen3 provided as a switch, does not help with these models. Directly against the API I also tried it in the system prompt and as a written-out instruction, with Qwen3.6 and Qwen3.8. The model reasoned every single time.

The only thing that works is the enable_thinking parameter. In n8n 2.39.9 the OpenAI Chat Model node has a field for it, Extra Body, which you add under Options via Add Option. Enter this:

{ "chat_template_kwargs": { "enable_thinking": false } }

With that, Qwen3.6 answers in just over a second with nothing but the category name. While you are at it, set Sampling Temperature to 0, since classification needs no creativity.

The reasoning mode has one more side effect. The model then puts two line breaks in front of its answer. The Basic LLM Chain trims them, the AI Agent does not. There the output was \n\nDer Rabatt beträgt 409,50 Euro. (the discount is 409.50 euros), and any exact text comparison further down fails. The AI Agent can use tools, by the way. With n8n’s Calculator tool, Qwen3.6 computed a discount correctly through a tool call, using 48 instead of 249 output tokens with Extra Body set.

Qwen3.8 comes with one more point. Its response times varied widely on the day of testing. In the afternoon my requests took 105 to 190 seconds, even when the answer was a single word. Just over half an hour later, the n8n runs took 5 to 8 seconds. The OpenAI Chat Model node gives up after 60 seconds by default. If you use Qwen3.8, set Timeout under Options to 300,000 milliseconds. How differently fast the two models are is covered in my re-measurement from August.

The HTTP Request node as an alternative

If your n8n version lacks the Extra Body field, or you want full control over the request, use the HTTP Request node. These settings worked in my test:

  • Method POST, URL https://inference.hetzner.com/api/v1/chat/completions
  • Authentication set to “Predefined Credential Type”, Credential Type “OpenAI”, then the same credentials as above
  • Send Body on, Body Content Type JSON, Specify Body “Using JSON”

This body structure worked, with the email coming from the previous node through an expression. I ran the test with a German system prompt; below is the English equivalent:

{
  "model": "Qwen/Qwen3.6-35B-A3B-FP8",
  "chat_template_kwargs": { "enable_thinking": false },
  "temperature": 0,
  "messages": [
    { "role": "system", "content": "Assign the email to exactly one category: Inquiry, Order, Complaint, Invoice, Other. Answer only with the category name." },
    { "role": "user", "content": {{ JSON.stringify($json.mail) }} }
  ]
}

JSON.stringify makes sure quotation marks and line breaks inside the email do not break the JSON body. The answer is then available at {{ $json.choices[0].message.content }}. In the test the node took 1.05 seconds for 4 output tokens.

Example workflow for sorting emails

A workflow you can rebuild in ten minutes sorts emails by what they are about. A Code node supplies five made-up German test emails, a complaint, an invoice question, a price inquiry, an order and an invitation to a local business association’s summer party. The Basic LLM Chain with the model node from above classifies each one. A Switch node then routes them with four rules of the form {{ $json.text }} is equal to “Complaint”. Everything else goes through the fallback output.

A hand sorting brown and light envelopes on a table Photo: cottonbro studio / Pexels

In the first run the prompt only listed the five category names. Four out of five emails landed correctly. The email saying the call-out fee had been charged twice on an invoice was classified as a complaint. That is not even unreasonable, but for accounting it is the wrong pile. With one line of explanation per category, all five were right. Translated from the German prompt I tested:

Assign the email to exactly one category.
Complaint: defects in delivered goods or in work carried out.
Invoice: questions or objections about an invoice, payment or reminder.
Inquiry: questions about prices, quotes or dates before an order.
Order: a binding order or commission.
Other: everything else.
Answer only with the category name.

One detail saves time when you build on this. The Basic LLM Chain only outputs the field text, the email body is gone afterwards. In the next node you get it back with {{ $('Mail').item.json.mail }}, where “Mail” is the name of the node that supplies the emails.

Later an Email Trigger (IMAP) replaces the Code node. Real customer emails, however, contain personal data, and for that Hetzner lacks the data processing agreement. So I recommend a simple approach. Build and tune the workflow with test emails on Hetzner. For live operation, point the credentials to a provider that offers such a contract.

The limit of ten requests per minute

According to the documentation, Hetzner allows 10 requests per key within 60 seconds. What I measured on 21 September looked different. Directly against the API, 67 of 70 sequential requests went through within 40 seconds. HTTP 429 errors mostly appeared when many requests arrived at the same time. I would not rely on that, since Hetzner can enforce the documented limit at any time. For n8n it matters anyway, because without further settings the HTTP Request node sends all items at once.

A wooden hourglass with yellow sand on a dark table Photo: Suki Lee / Pexels

Setup in n8n Items Result Duration
HTTP Request without batching 12 all successful 1.2 s
HTTP Request without batching 40 17 with HTTP 429 1.4 s
same with Retry On Fail, 3 tries 40 workflow stops 6.2 s
HTTP Request, batching 1 item every 6.5 s 12 all successful 72.5 s
Basic LLM Chain, Batch Size 2, delay 13 s 12 all successful 68.3 s

The n8n documentation on rate limits names two approaches, Retry On Fail and Batching. Retry On Fail is the obvious switch and does not help here. It repeats the whole node with all 40 items, so the same burst hits the limit three times. n8n also caps the wait between tries at 5 seconds. If On Error is set to “Continue”, n8n repeats nothing at all, and the errors simply end up as items in the output.

What works is a steady pace. The HTTP Request node has a Batching option under Options. With Items per Batch 1 and Batch Interval (ms) 6500 you stay safely below the documented limit. In the Basic LLM Chain the counterpart is called Batch Processing, with Batch Size and Delay Between Batches. There is a trap here. With a Batch Size of 1 the node ignores the delay entirely, and 12 items ran through in 6 seconds. The delay only kicks in from a Batch Size of 2, and with 13 seconds between batches you are back at 10 requests per minute. For a nightly run over a few hundred items, that pace is perfectly fine.

What this is good for

Hetzner Inference in n8n fits everything that contains no personal data and is not time-critical. Typical examples are product descriptions drafted from a spreadsheet, summaries of internal notes, tags for your own documents, or a workflow you tune with test data until only the credentials change for live operation. For tasks like these, Qwen3.6 with reasoning switched off is fast enough and costs nothing.

It does not fit real customer data, because the contract is missing. Workflows that have to run reliably every day do not belong there either. Hetzner itself calls the service an experiment. Since the launch in July, models and limits have changed several times. A workflow that distributes incoming requests to your team every morning needs a foundation with commitments. Because everything in n8n hangs on the credentials, moving there later takes minutes.

If you want to set up an AI workflow in n8n and need to work out which provider fits your data, feel free to get in touch. I will go through the workflow with you and tell you which model and which settings suit it.

FAQ

Can I use Hetzner Inference in n8n for free?+

Yes. According to Hetzner, the Inference API is free of charge as long as it has experimental status. You need a key from the experiments area of your Hetzner account and create credentials of type OpenAI in n8n with the base URL https://inference.hetzner.com/api/v1. There is no availability guarantee, and Hetzner does not offer a data processing agreement for the experiment.

Why does the OpenAI node in n8n return a 404 error with Hetzner?+

Because the Use Responses API toggle in the OpenAI Chat Model node is switched on by default, at least in n8n 2.39.9. n8n then calls OpenAI's Responses interface, which Hetzner does not provide. The error reads The resource you are requesting could not be found and sounds like a wrong model name. Switch the toggle off and the request goes through Chat Completions and works.

How do I turn off Qwen's reasoning mode in n8n?+

In the OpenAI Chat Model node, add the Extra Body option under Options and enter { "chat_template_kwargs": { "enable_thinking": false } }. Alternatively, send the request with the HTTP Request node and put the parameter straight into the JSON body. Adding /no_think to the prompt or an instruction in the system prompt had no effect in my test.

Can I process customer emails with Hetzner Inference in n8n?+

Not with real customer data under GDPR, because Hetzner does not offer a data processing agreement for the experiment. The fact that Hetzner says it does not store request content does not replace that contract. For a prototype with made-up test emails it works well. For live operation you later swap the credentials for a provider with a contract, and the workflow itself stays the same.

Want to know more?

In a free intro call we discuss how you can use these topics for your company. Not a sales pitch, but an honest assessment.

Book a free intro call