Skip to content
HomeHome
DE
WhatsAppMailPhone
← All articles
Gemini 3.7 Flash and 3.8 Flash cost 0.13 to 0.83 cents per office task
KI

Gemini 3.7 Flash and 3.8 Flash cost 0.13 to 0.83 cents per office task

Photo: SHVETS production / Pexels

I measured what four typical office tasks cost with Gemini 3.7 Flash and 3.8 Flash. At the introductory price it is 0.13 to 0.83 US cents per task, from 1 January 2027 twice that. Most of the cost is the model's thinking. That is also where the savings are.

Eric MengeAuthorEric MengeOwner & web developer at EMIT Solution
Published
Reading timeca. 8 min

In short

  • A typical office task costs between 0.13 US cents (sorting an email into a category) and 0.83 US cents (summarising a document of almost 1,500 words) with Gemini 3.7 Flash or 3.8 Flash at the introductory price, measured on 21 September 2026 with default settings.
  • On 1 January 2027 both models double their input and output prices to 1.50 and 7.50 US dollars per million tokens. The cost of every single task doubles with them.
  • With default settings, 46 to 89 percent of the cost goes on thinking tokens, which Google bills at the output rate. When sorting an email, the answer is three tokens long and the thinking before it around 300.
  • With the thinking level set to low, the cost of the four tasks on Gemini 3.8 Flash fell by 64 percent with no visible loss of quality. From 2027, 3.8 Flash on low is therefore cheaper than 3.7 Flash on default settings today.

The introductory price for Gemini 3.7 Flash and Gemini 3.8 Flash ends on 31 December 2026. From 1 January 2027 Google charges 1.50 instead of 0.75 US dollars per million input tokens for both models and 7.50 instead of 3.75 US dollars per million output tokens (as of 21 September 2026, according to Google’s pricing page). A token is the billing unit, a chunk of text that is often shorter than a word. My German test document of 1,459 words came to 2,251 tokens including the instruction.

The token price alone does not tell you what a single task costs. I use Gemini 3.7 Flash in my own applications, so the change at the turn of the year affects me directly. On 21 September I therefore measured four tasks that small businesses deal with every day, with both models and at both price levels.

Prices and discounts at a glance

Gemini 3.7 Flash has been generally available since 13 August 2026, Gemini 3.8 Flash since 2 September. Both cost exactly the same per token and the introductory price ends on the same day for both. One note on the pricing page matters for every calculation. The output price includes thinking tokens. That is the text the model generates internally before the actual answer in order to work through the task. You never see it, but you pay for it like answer text.

There are two discounts. Through the Batch API, where Google processes collected requests within up to 24 hours, everything costs half. The Flex tier has the same price and answers within minutes, but without firm guarantees. From 2027 a batch call will therefore cost as much as a normal call does today.

The second discount is caching. Repeated input is then billed at a tenth of the input price. According to the caching documentation, however, automatic caching only kicks in from 4,096 input tokens on 3.7 and 3.8. Explicitly created caches add an hourly storage fee. None of my four tasks came close to the threshold. For short office tasks the discount is irrelevant. It only pays off when every call starts with the same long instructions or documents.

For audio, 3.7 and 3.8 list a single input price with no separate audio rate. Gemini 3.1 Flash-Lite, by contrast, charges twice its text price for audio and 2.5 Flash-Lite three times. The two Flash-Lite models I measured for comparison have no expiring introductory price. Gemini 3.5 Flash-Lite costs 0.30 and 2.50 US dollars per million tokens, 2.5 Flash-Lite 0.10 and 0.40 US dollars.

How I measured

I wrote the texts myself, in German, for a fictional joinery business. The four tasks:

  • reply to a customer enquiry of around 150 words without quoting prices and with a proposed date for an on-site measurement
  • extract number, date, line items, amounts and due date from an invoice as JSON, a format software can process directly, enforced by a response schema with fixed fields
  • sort a customer email into one of five categories
  • summarise an internal proposal of almost 1,500 words in five sentences

Each task ran five times per model through the standard API, with no extra parameters, so with whatever thinking the model applies on its own. For 3.7 and 3.8 I added five runs each at the lowest and the highest thinking level, 160 calls in total. From every response I took the token counts for input, answer and thinking and calculated the cost from them. The whole measurement cost around 0.54 US dollars.

I checked quality as well. All 40 JSON outputs were valid and correct, down to the correctly calculated due date. All 40 classifications said complaint, even though the email also mentions the unpaid invoice. None of the reply emails quoted a price.

Front panel of a server with glowing blue status lights Photo: panumas nikhomkhai / Pexels

What a task costs

Cost per 1,000 tasks in US dollars, each with default settings:

Task 3.7 Flash until 31 Dec 2026 3.7 Flash from 1 Jan 2027 3.8 Flash until 31 Dec 2026 3.8 Flash from 1 Jan 2027 3.5 Flash-Lite
Reply to customer enquiry 4.00 7.99 3.85 7.70 0.59
Invoice data as JSON 3.01 6.01 4.26 8.51 1.35
Sort email into category 1.27 2.54 1.39 2.79 0.07
Summarise document 7.38 14.75 8.28 16.56 1.22
All four together 15.65 31.30 17.78 35.55 3.23

Per task that is between 0.13 US cents for sorting an email with 3.7 Flash and 0.83 US cents for the summary with 3.8 Flash at the introductory price. From January every figure doubles, because input and output prices rise in the same proportion. A business that has 2,000 incoming emails a month sorted and answered pays around 10.50 US dollars for it with 3.7 Flash today and around 21 US dollars from January.

That is still a modest amount. But anyone who draws up an annual budget today, or quotes a fixed price for an AI application, and plugs in the introductory price is working with the wrong figure.

Desk calendar showing December up to the 31st Photo: Matheus Bertelli / Pexels

Most of what you pay for is thinking

What stands out is where the money goes. With default settings, 46 to 89 percent of the cost goes on thinking tokens, depending on the task. Sorting is the clearest case. The answer is a single word, three tokens. Before it, 3.7 Flash thinks for 297 tokens on average and 3.8 Flash for 329. For the summary, around 1,300 to 1,500 thinking tokens stand against an answer of around 200 tokens.

Thinking also costs time. The summary took just over six seconds (median) with default settings. In the runs where the model did not think, it took under two.

3.8 Flash thinks more and the premium depends on the task

Comparison articles quote a figure of 45 percent higher costs for 3.8 Flash. It goes back to Artificial Analysis, where one task of their Intelligence Index cost 0.58 instead of 0.40 US dollars at high thinking effort, with around 48,000 output tokens per task. Google itself writes on its model page that 3.8 deliberately uses more tokens on longer and more complex tasks. For everyday tasks Google recommends a lower thinking level.

My tasks are orders of magnitude smaller. On default settings, 3.8 cost just under 14 percent more than 3.7 across all four. The range is wide. The reply email was actually 4 percent cheaper with 3.8, sorting 9 percent and the summary 12 percent more expensive. Only on the invoice did 3.8 come close to the quoted figure, at 42 percent, because it thought almost twice as long there, 675 instead of 371 thinking tokens on average. At the high thinking level the two models were closer together, with 3.8 costing 9 percent more overall.

I could not see a quality difference between the two models on these tasks. Both passed the same checks in every run and both made the same mistake with the proposed dates, more on that below.

The low thinking level is the biggest lever

On current models, thinking is controlled with a level instead of a token budget. 3.7 and 3.8 accept low, medium and high, with medium as the default, as described in the thinking documentation. The level is set in the request:

"generationConfig": {
  "thinkingConfig": { "thinkingLevel": "low" }
}

The effect in average thinking tokens, with the cost of 1,000 runs of all four tasks at the 2027 price underneath:

3.7 default 3.7 low 3.8 default 3.8 low
Reply to customer enquiry 829 808 775 157
Invoice data as JSON 371 218 675 0
Sort email into category 297 116 329 143
Summarise document 1,321 0 1,541 0
Cost in US dollars for 1,000 runs of all four tasks from 2027 31.30 19.36 35.55 12.97

On 3.8 Flash, low had a sweeping effect. For the invoice and the summary the model stopped thinking altogether, for the reply email in four of five runs. Across all four tasks the cost fell by 64 percent. 3.7 Flash follows the level less consistently. For the summary thinking disappeared completely, for the invoice and sorting it dropped by 40 to 60 percent. For the reply email it stayed practically unchanged. Overall 3.7 on low saved 38 percent.

That turns the ranking around. On low, 3.8 Flash is a third cheaper than 3.7 Flash on low. From January 2027 it even costs less than 3.7 Flash on default settings today, 12.97 against 15.65 US dollars per 1,000 runs of all four tasks. I saw no loss of quality. JSON, category and sentence count were correct in every run.

There are two pitfalls. Both 3.7 and 3.8 reject the level minimal, which other models in the family accept, with an error. And both accept the older parameter thinkingBudget set to 0 without an error message but keep thinking anyway, for 103 and 112 tokens in my test. If you rely on it, you pay for thinking you believe is switched off.

Flash-Lite costs a fraction

Gemini 3.5 Flash-Lite works with minimal thinking by default and did not bill a single thinking token in any run. For all four tasks together it came to 3.23 US dollars per 1,000 runs, around a fifth of 3.7 Flash today and a tenth from January. The older 2.5 Flash-Lite came in at 0.53 US dollars. Both answered in under two seconds (median). Gemini 3.1 Flash-Lite would be slightly cheaper per token than 3.5, but according to Google it will be shut down on 7 May 2027.

For extraction and sorting every result was correct. The summary showed the difference. In three of five runs, 2.5 Flash-Lite wrote that the changeover described in the document takes three months, when the document says six. 3.5 Flash-Lite stayed factually correct but made one grammatical mistake. For tasks where a schema keeps the output in check, a small model is enough. For texts a person reads and is expected to trust, I would stay with the Flash model.

A side finding in the reply emails

In 20 of 30 reply emails from 3.7 and 3.8, the model proposed a concrete appointment with a weekday, for example Tuesday, 22 October. In all 20 cases the weekday did not match October 2026. Seven of the dates would have fallen on a Saturday. The model does not know today’s date unless it is in the request.

The cross-check was clear. With the sentence “Today is Monday, 21 September 2026” (in German) at the start of the system instruction, all twelve weekdays mentioned in ten further runs were correct. The sentence costs 16 tokens per call.

A pile of euro coins on a dark surface Photo: Panos and Marenia Stavrinos / Pexels

What you can do now

Any budget that runs past the turn of the year should use the 2027 price, twice what the bill shows today. The thinking level should be set per task and not left to the model. For sorting, extraction and summaries, low was enough in my runs. That is exactly where the lever is biggest. Anything that does not need to be finished immediately, such as processing incoming invoices overnight, can run through Batch at half price. If you sort or extract large volumes, Flash-Lite is at least worth a test.

For me the most important finding is that the thinking level moves more than the choice between 3.7 and 3.8. The models were 14 percent apart on default settings, while default and low were up to 64 percent apart. If you are reviewing your setup before December anyway, test the combination of 3.8 Flash and the low level against your own tasks.

One point applies regardless of the model. If you make an application with Gemini behind it available to users in the European Economic Area, Switzerland or the UK, Google’s Gemini API Additional Terms require a paid tier anyway. The free tier is meant for your own testing. How long providers log requests is covered in my article on zero data retention. If you want to experiment without paying anything, Hetzner’s free inference API is a playground, though without any commitment for production use.

If you want to know what the AI tasks in your business will cost from January and which thinking level is enough for them, feel free to get in touch.

FAQ

How much will Gemini 3.7 Flash cost from 2027?+

From 1 January 2027 Gemini 3.7 Flash costs 1.50 US dollars per million input tokens and 7.50 US dollars per million output tokens, thinking tokens included. Until 31 December 2026 an introductory price of 0.75 and 3.75 US dollars applies. Gemini 3.8 Flash has exactly the same prices and the same cut-off date (as of 21 September 2026).

Is Gemini 3.8 Flash more expensive than 3.7 Flash?+

Not per token, but usually per task, because 3.8 thinks more. In my measurement of four office tasks, 3.8 on default settings cost just under 14 percent more in total, 42 percent more for invoice extraction and 4 percent less for the reply email. With the thinking level set to low, 3.8 was a third cheaper than 3.7 on low.

How do I turn down thinking on Gemini 3.7 Flash and 3.8 Flash?+

With the thinkingLevel parameter inside the thinkingConfig of the request. Both models accept low, medium and high. Medium is the default. Both reject the level minimal with an error. The older parameter thinkingBudget set to 0 is accepted without complaint, but the models keep thinking anyway.

Is Gemini Flash-Lite good enough for office tasks?+

For sorting and extraction with a fixed response schema, every result in my measurement was correct at a fraction of the cost. Gemini 3.5 Flash-Lite cost 3.23 US dollars per 1,000 runs of all four tasks, against 15.65 US dollars for 3.7 Flash today. For summaries, the older 2.5 Flash-Lite made a factual error in three of five runs, so the larger model is the safer choice there.

Want to know more?

In a free intro call we discuss how you can use these topics for your company. Not a sales pitch, but an honest assessment.

Book a free intro call