Skip to content
HomeHome
DE
WhatsAppMailPhone
← All articles
AI transcription invents entire voice messages from silence
KI

AI transcription invents entire voice messages from silence

Photo: Werner Pfennig / Pexels

I had five Gemini models transcribe empty recordings. With a short company description in the system prompt, 148 of 150 runs returned invented text, often with a caller name, a request and a phone number. What helps if you transcribe voice messages with AI.

Eric MengeAuthorEric MengeOwner & web developer at EMIT Solution
Published
Reading timeca. 8 min

In short

  • On 21 September 2026, five Gemini models returned invented text for digital silence and faint noise in 148 of 150 runs once the system prompt described a company. Not a single response was empty.
  • The inventions draw on the context. 130 of the 148 invented transcripts mentioned products or people of the described company. 65 contained a made-up phone number.
  • Gemini 3.5 Flash, 3.5 Flash-Lite and 2.5 Flash followed an instruction to return only a placeholder for silence in 88 of 90 cases. Gemini 3.7 Flash and 3.8 Flash still produced invented text in 46 of 60 cases despite that instruction.
  • A level check before the model reliably holds back empty recordings. In the test, silence measured minus infinity, faint noise -49.5 dBFS and the spoken test message -13.2 dBFS, so a threshold of -40 dBFS separates them cleanly.

Anyone who has voice messages, voicemail or meetings transcribed automatically expects an empty transcript when the recording is empty. On 21 September 2026 I tested whether that holds for the current Gemini models. It does not. With a short company description in the system prompt, pure silence came back as text in 148 of 150 runs, usually as a complete customer message with a name, a request and a callback number.

For OpenAI’s Whisper the phenomenon has been documented since 2024. A study led by information scientist Allison Koenecke at Cornell University found entire sentences that did not exist in the audio in roughly one percent of the transcripts examined. Speakers with long pauses were affected most. I could not find a comparable test of the current Gemini models, so I measured it myself.

How I tested

I generated six test files, all 16 kHz mono WAV. Five contain digital silence, meaning nothing but zeros, at lengths of 0.5, 1, 1.5, 3 and 10 seconds. The sixth contains three seconds of very faint, even noise at -50 dBFS (decibels relative to full scale, where 0 is the technical maximum). As a control I added a spoken message of 13.5 seconds, created with the German text-to-speech voice built into Windows. That shows whether the models still transcribe normal messages correctly.

I tested five models through the Gemini API: Gemini 3.8 Flash, 3.7 Flash, 3.5 Flash, 3.5 Flash-Lite and 2.5 Flash. Every file ran on every model under three conditions, five times each. All prompts were in German, the test was about German voice messages:

  • A: only the instruction “Transcribe the audio recording verbatim in German.”
  • B: the same instruction plus a system prompt describing a fictional carpentry workshop, including products such as dining tables and built-in wardrobes and typical requests such as delivery dates or complaints.
  • C: like B, with one added sentence: “If nothing spoken can be heard, reply only with [keine Sprache].”

Condition B is the realistic case. Anyone building transcription for a business voicemail gives the model context so that product terms and names are spelled correctly. I set the models’ thinking effort to low. A preliminary run with default settings showed the same picture, all five models turned one second of silence into a message. In total there were 525 calls, which cost 0.40 US dollars altogether.

Man recording a voice message on his smartphone Photo: Theo Decker / Pexels

With company context every model invents almost every time

The table counts how often each model returned invented text. Anything containing at least one word that was never spoken counts. Pure timestamps such as “00:00” or sound tags such as “(Rauschen)”, German for noise, do not. Each cell covers 30 runs, six silent or noisy files times five repetitions.

Model A: instruction only B: with company description C: B plus placeholder rule
Gemini 3.8 Flash 28 29 24
Gemini 3.7 Flash 27 30 22
Gemini 3.5 Flash 4 29 2
Gemini 3.5 Flash-Lite 0 30 0
Gemini 2.5 Flash 17 30 0
Total 76 of 150 148 of 150 48 of 150

The middle column is the real finding. No model returned an empty response even once under condition B. 141 of the 150 responses were complete sentences with a median of 51 words. 130 of the 148 invented transcripts mentioned products or people of the described workshop. From one second of silence, Gemini 3.7 Flash produced this message, translated here from German.

“Good day, this is Karin Becker. I wanted to ask how our walnut dining table is coming along. You said it would probably be ready by the middle of the month. Can you give me a firm delivery date yet?”

Other runs reported a creaking fourth stair, a sticking wardrobe door or a chipped table edge. 65 of the invented transcripts contained a phone number, 64 an explicit request for a callback. The length of the recording made no difference. Half a second of silence produced a complete customer message just as reliably as ten seconds.

Carpentry workbench with tools, a steel ruler and wood shavings Photo: Tima Miroshnichenko / Pexels

Without company context the picture is more mixed. Gemini 3.7 Flash and 3.8 Flash still invented text in 55 of 60 cases, just without a theme, things like “Shall we go to the cinema tonight?” or a greeting from a car dealership that does not exist. Gemini 3.5 Flash almost always wrote only “00:00” or its spoken form, Gemini 3.5 Flash-Lite mostly returned nothing. Gemini 2.5 Flash returned its own instruction as the transcript eleven times, mostly word for word. That fits the pattern too, because the instruction was the only text the model had.

The placeholder rule only helps with some models

The obvious reaction is one more line in the prompt, which is exactly what condition C tested. With Gemini 3.5 Flash, 3.5 Flash-Lite and 2.5 Flash it works almost completely, 88 of 90 responses were the correct placeholder. Gemini 3.7 Flash and 3.8 Flash, on the other hand, complied in only 13 of 60 cases and kept writing messages from invented customers in the rest.

The control with the spoken message ran cleanly under all three conditions. All 75 transcripts were correct. No model returned the placeholder for spoken text even once. So the rule does no harm, it just does not protect with every model.

Why this happens

A language model does not listen the way a person does. It produces the answer most likely to follow the entire input. Asked to transcribe, that answer is a transcript, even when there is nothing to transcribe. If the audio carries no signal, the rest of the context decides what the transcript looks like. A system prompt describing a carpentry workshop yields a carpentry message. With only the instruction present, the model may simply copy the instruction.

That leads to an uncomfortable rule. The better the system prompt describes the business, the more convincing the invention. The very context that gets names and product terms right in real messages turns an empty recording into a credible customer request.

Black office telephone on a desk next to a document tray Photo: Ron Lach / Pexels

Why this is dangerous in practice

If you dictate yourself, you see the transcript immediately and notice when it is wrong. Incoming voice messages are different. There the transcript is often the only thing anyone reads, because that is the whole point of having it. Hardly anyone listens again to a message whose text looks complete and plausible.

Empty recordings come about in very ordinary ways. A caller hangs up silently after the beep, a phone records in a jacket pocket, or a microphone muted in Windows delivers no error message, just silence. If such a recording lands in a setup like condition B, a ticket appears in the system that looks like a genuine complaint. In the harmless case the office calls a number that does not exist. In the awkward case the number belongs to someone who never called.

What helps

Check the level before the model sees the file

The most effective measure is not to send empty recordings in the first place. A level measurement is enough, meaning the question of how loud the loudest part of the recording is. This Python function measures the average loudness (RMS) of every 50 millisecond segment and returns the loudest, in dBFS:

import wave, numpy as np

def loudest_segment_dbfs(path, window_ms=50):
    with wave.open(path, "rb") as w:
        rate = w.getframerate()
        x = np.frombuffer(w.readframes(w.getnframes()), dtype="<i2") / 32768.0
    n = int(rate * window_ms / 1000)
    frames = x[: len(x) // n * n].reshape(-1, n)
    rms = np.sqrt((frames ** 2).mean(axis=1)).max()
    return 20 * np.log10(max(rms, 1e-10))

if loudest_segment_dbfs("message.wav") < -40:
    print("No speech detected, recording will not be transcribed.")

On my test files the silence measured minus infinity, the noise -49.5 dBFS and the spoken message -13.2 dBFS. A threshold of -40 dBFS separates them with plenty of margin. You should still verify the exact value against a few real recordings from your phone system or microphone, because every line has its own background noise. Voice messages from messenger apps usually arrive as Opus files, which you convert to WAV first with ffmpeg -i message.opus -ar 16000 -ac 1 message.wav. After exactly that round trip through Opus, the test message measured -13.8 dBFS and the silence stayed at minus infinity.

Screen showing an audio waveform in an editing program Photo: Egor Komarov / Pexels

A minimum length only as a pre-filter

A minimum length sounds obvious but does not protect on its own. Under condition B, half a second of silence produced a complete sentence in 25 of 25 runs. Ten seconds of silence did exactly the same. A lower limit of about one second catches accidental taps, nothing more.

Label uncertain transcripts

The placeholder rule still belongs in the prompt, because it works reliably with three of the five models and does not affect real speech. When the placeholder comes back, the interface should show that clearly instead of creating an empty message. It is also worth labelling every automatic transcript as such and keeping the original recording one click away. Anyone in doubt can then listen briefly.

Transcribe twice and compare

The most reliable check after the fact is repetition. The spoken test message produced almost identical text twice with every model. The 150 compared pairs agreed by at least 98 percent. The differences were trivia such as “10 Uhr” against “10:00 Uhr”. Invented transcripts agreed by only 17 percent at the median, the most similar pair reached 92 percent. The model tells a different story on every run. That is how you recognise the invention.

Two caveats apply. My control was a clean synthetic voice. With noisy phone recordings, genuine transcripts also differ more from each other, so the threshold has to be lower. At 70 percent, 4 of 399 invented pairs would have slipped through in my test. The comparison also misses outputs that are identical every time, such as the copied instruction or “00:00”. It complements the level check, it does not replace it.

The second call costs little. In the test, Gemini 3.7 Flash billed 25 tokens per second of audio. At 0.75 US dollars per million input tokens that is about 0.11 US cents of input cost per minute of recording (as of 21 September 2026, introductory price until the end of 2026 according to Google’s price list). I broke down what the Flash models cost per task overall in a separate cost comparison.

What to take away

AI transcription has no state called “heard nothing” unless you give it one. With company context, all five Gemini models tested turned silence into credible customer messages. The two newest ones could hardly be talked out of it even with a clear instruction. The solution therefore does not lie in the model but in front of it. A level check of a dozen lines of code, visible labelling and, for important workflows, a second pass are enough to make sure an empty recording no longer becomes a ticket.

If you want voice messages, voicemail or meetings transcribed automatically and need to be sure that only genuine messages reach your inbox, get in touch.

FAQ

Why does AI transcription invent text when nothing is said?+

A language model returns whatever answer seems most likely given the whole input. Asked to transcribe a recording, a transcript is more likely than nothing at all. If the audio carries no signal, the model fills the gap with whatever fits the rest of the context. In my test, with a carpentry workshop described in the system prompt, that almost always meant messages from customers of that workshop.

Which Gemini model is most reliable with empty recordings?+

With the instruction to write only [keine Sprache] (German for no speech) when nothing is audible, Gemini 3.5 Flash-Lite and 2.5 Flash returned the correct placeholder in 30 of 30 runs in my test, Gemini 3.5 Flash in 28 of 30. Gemini 3.7 Flash and 3.8 Flash complied in only 13 of 60 cases. Without that instruction and with company context, all five models almost always invented text. Choosing a model is therefore not enough, the check belongs in front of the model.

Does it help to tell the AI in the prompt not to make things up?+

For some models yes, for the newest ones hardly. The placeholder rule cut invented transcripts in my test from 148 to 48 out of 150. 46 of the remaining 48 came from Gemini 3.7 Flash and 3.8 Flash. An instruction is an additional safety net, not a replacement for a technical check of the audio before the call.

How can I spot an invented transcript after the fact?+

The simplest way is to transcribe the same recording twice and compare the texts. A spoken test message produced almost identical text twice in my test, with at least 98 percent agreement. Invented transcripts agreed by only 17 percent at the median, because the model tells a different story each time. It also helps to keep the original recording one click away from every transcript.

Want to know more?

In a free intro call we discuss how you can use these topics for your company. Not a sales pitch, but an honest assessment.

Book a free intro call