Does Transcription Work on Music-Heavy Reels?
success:true with no warning field.How we tested this
We took one 10.2-second clip of speech whose text we know word for word (29 words), and mixed it under music beds we synthesised ourselves — a sustained three-note pad, a drum-style bed, a bass-heavy bed, and a dense mix of broadband noise plus chord tones modulated at about two beats per second. Every file went to our own /api/transcribe endpoint, which runs @cf/openai/whisper on Cloudflare Workers AI. Levels were measured rather than guessed: the speech averages -18.1 dB, the dense bed -17.7 dB at its base level, so "1x" in the table below means the music and the voice are roughly equal.
Two honest limits: the music is synthesised, not licensed tracks, and it is one voice reading one sentence. Treat the shape of these numbers as the finding, not as a benchmark you can quote for your own audio.
Steady music is much less of a problem than you would expect
The tonal and percussive beds barely registered. Across the pad bed at 0.25x, 0.5x, 1x, 2x and 4x the voice level, all 29 words came back every time. The bass-heavy bed did the same at 0.5x through 4x. The drum bed held 29 of 29 up to 2x and lost a single word at 4x. Music that occupies a few narrow bands of the spectrum is, to the model, mostly not speech-shaped, so it gets ignored.
This is the part most advice gets wrong. People assume a loud beat is the enemy. In our runs, a beat you can clearly hear under a clear voice changed nothing.
The dense mix is where it falls apart
Once the bed carries energy across the whole spectrum and keeps changing, the level starts to matter a great deal. We scored every result by aligning the output against the known 29 words and counting how many survived.
- Dense bed 6 dB below the voice: 29 of 29 words.
- Dense bed at equal level: 28 of 29 — "brown" came back as "ground".
- Dense bed 6 dB above the voice: 24 of 29 — "Instagram transcript tool" became "Instagram Transfer Tool", and "quick brown fox" became "quick round box".
- Dense bed 12 dB above the voice: 15 of 29, in a 32-word response that reads like a sentence and shares almost nothing with what was said.
- Voice dropped to a quarter of the music level: 16 of 29, again as fluent but wrong text.
- Voice dropped to a tenth of the music level: the entire response was two characters.

Three failure shapes, and none of them is an error
The first shape is quiet substitution: a handful of words change and the rest is right, which is dangerous because the transcript still looks trustworthy. The second is collapse — files containing nothing but music came back as a single token, and a voice buried at a tenth of the music level came back as "I'm". Nothing in the pipeline treats that as failure.
The third shape is the one to watch for. When we added a second voice standing in for sung lyrics at 6 dB above the target voice, the output grew to 164 words and only 8 of our 29 reference words appeared: the model had transcribed the wrong track and produced a long, grammatical, completely fictional paragraph. A response to /api/transcribe carries exactly six fields — success, transcript, language, engine, fileSize and fileName. There is no confidence score, no per-word probability, and no flag for "this audio was hard", so nothing in the response distinguishes a clean run from an invented one.
Nothing you do to the file afterwards fixes it
We tried the usual repairs on the worst file. High-pass filtering at 100 Hz to remove the bass returned a byte-identical transcript. Loudness normalisation also returned a byte-identical transcript. Noise reduction made it worse, dropping from 15 of 29 to 10 of 29. Downsampling to 8 kHz was rejected outright with HTTP 500.
Two tricks that work elsewhere did not work here either. Re-uploading the same file three times produced byte-identical output each time, including the errors — so the "run it twice and diff the two transcripts" method that catches wrong words in noisy audio catches nothing when the cause is music. Cutting the damaged file into four-second chunks gave 10, 6 and 4 of 29 words: smaller pieces did not isolate the good parts, because the music runs under all of them.
What actually helps
Sanity-check the length against the audio. Our clean 10.2-second clip is 29 words. If a 60-second Reel comes back with three words, or with a paragraph that has nothing to do with the subject, the audio lost and the model filled the gap.
- Cut the music-only stretches off before uploading. A six-second music-only intro followed by real speech came back 29 of 29, with no invented padding, so trimming matters more than filtering.
- If you made the Reel yourself, transcribe the camera audio from before the music bed was mixed in. Once the two are summed into one track, no filter we tested could pull them apart again.
- Do not trust fluent text that came out of a section where you cannot hear a voice. Every one of our bad results was grammatical English.
- Treat a one-word or two-word transcript as a failure, not as a short answer.
What we cannot tell you
We do not publish an accuracy percentage, because one clip and one voice is not a test set and a single number would pretend otherwise. Real licensed music with sung vocals may behave differently from our synthesised beds — we do not have the rights to test with commercial tracks, and we are not going to guess at what they would do. What we can say is measured: the failure is level-dependent, it is silent, and it cannot be repaired after the file exists.
Got a Reel you want as text? Upload the file and download the transcript as TXT, SRT, or VTT →
FAQ
Does background music make an Instagram transcript inaccurate?
Only when the music is dense and broadband, and only once it is loud relative to the voice. A steady pad or a beat at up to 12 dB above the voice left all 29 words of our reference clip intact. A dense mix at the same level started changing words, and at 12 dB above the voice replaced half the sentence with invented text.
Why does the tool not warn me that the audio was difficult?
Because the response has no field for it. Our endpoint returns six fields: success, transcript, language, engine, fileSize and fileName. The model returns its best guess and no confidence score, so a confident hallucination and a clean transcript look identical in the response.
Can I strip the music out of a Reel before transcribing it?
Not with the filters we tried. High-pass filtering at 100 Hz and loudness normalisation both returned byte-identical transcripts, noise reduction made the result worse, and downsampling to 8 kHz was rejected with HTTP 500. Music and speech occupy the same frequencies, so there is no filter that separates them once they are mixed.
Do music-only parts of a Reel produce empty transcripts?
No, and that is the problem. Files containing only music came back as a single token rather than an empty string, and a voice buried at a tenth of the music level came back as "I'm". Garbage arrives looking like a short answer, not like a failure.
Should I split a music-heavy Reel into chunks?
It did not help here. Four-second chunks of a music-damaged file returned 10, 6 and 4 of the 29 reference words. Chunking does rescue some wrong-word problems, but when the damage is continuous music under the whole clip, every chunk carries the same damage.