← Instagram Transcript

Instagram Transcript for Podcasters: Reels Into Show Notes

2026-10-10
Short answer

Yes. Export the audio, cut it into five-minute chunks, transcribe each one, and paste the text into your show-notes draft. What comes back is the words and only the words: across the 19 uploads we ran today, every transcript arrived with zero timestamps, zero speaker labels and zero paragraph breaks. The words are the raw material; the timestamps, the chapter markers and the structure are still yours to add. Below is exactly what we measured, and the chunk sizes that actually get through.

What a podcaster actually gets back

The response is the same shape for a 15-second Reel and for a five-minute chunk of an episode. Six fields, and the model is named in one of them so you can check it independently:

What we measured today: 19 uploads

All of it ran this morning against our own /api/transcribe endpoint, timed end to end with the upload included. The speech is one short clip looped to length, so treat the word counts as a shape, not as a natural-speaking-rate benchmark.

Seven measured results from 19 uploads: Reel clips ten of ten, five-minute chunks four of four, a twenty-minute episode in 196.8 seconds, and two HTTP 500s with different causes
19 uploads to our own endpoint today. The two failures have different causes: one is size, one is length.

Two separate ceilings decide your chunk size

They look like one problem and they are not, which matters because the fix for each is different:

A workflow that fits a real episode

This is the sequence that survived every ceiling we hit today:

The parts of show notes this tool does not produce

Stated plainly, because finding this out after an hour of transcription is annoying:

Where the words need checking before you publish

Show notes go out under your name, so the cost of a wrong word is higher than in a private draft:

What this tool does with your episode file

You are uploading unpublished audio, so this is worth stating exactly. The file leaves your machine: it is POSTed to /api/transcribe over an encrypted connection, decoded into memory, handed to the model once, and discarded when the response is sent. There is no account, no stored copy, no training use, and no result you can come back to later — a GET to the same endpoint returns 404. If your episode is embargoed, the practical caution is the same as with any third-party processor: transcribe a chunk you are comfortable sending, or wait until it is public.

Affiliate link

Chunking an episode only helps if each chunk is in a container this tool will read. Wondershare UniConverter converts the export to M4A or MP3 on your own machine, and writes the pieces out at a bitrate small enough to clear the upload ceiling.

FAQ

Can I transcribe a whole podcast episode in one upload?

Not reliably. Two ceilings stop you: a size ceiling between 3.18 MB and 3.35 MB, which refuses in a couple of seconds, and a length ceiling that failed a 10-minute chunk after about two minutes even though the file was only 2.4 MB. Five-minute chunks at 64 kbps mono cleared both in every run we did today — four out of four, averaging 49.2 seconds each.

Does the transcript come with timestamps I can use as chapter markers?

No. The response has six fields and none of them carry timing. Across today's successful uploads we counted zero timestamps and zero line breaks in the text. Our SRT and VTT download builds a single cue with zero duration, because dividing the text evenly would mean inventing numbers. For real chapter marks, align the audio in a subtitle editor, or use your chunk filenames as approximate offsets.

Will it tell me which host is speaking?

No. There is no speaker separation and no option to request it. Two voices on one track come back as a single run of text. Overlapping speech made this worse in our earlier tests: repeated phrases and dropped words, with the response still reporting success.

Is my unpublished episode file stored anywhere?

No, but be precise about what that means: the file does leave your browser. It is POSTed to /api/transcribe over TLS, decoded into memory, processed once by the model, and discarded as soon as the response is sent. There is no account, no saved copy, and no training use, and there is no result page to come back to — a GET to the same endpoint returns 404. Copy the text before you close the tab.

Why can't I just paste the Reel URL?

Because Instagram blocks server-side fetching of its media. Every route we tested — direct fetch, the oEmbed endpoint without a token, and the unauthenticated GraphQL path — returns nothing usable, which is why this site has no URL box. Tools that do accept a link are usually asking you to log in first, and handing over your Instagram password to a transcript tool is a trade we are not willing to make.

How long does a 60-minute episode take?

About 10 minutes of waiting, if you cut it into twelve five-minute chunks. That is arithmetic from today's mean of 49.2 seconds per chunk, not a timed 60-minute run, and the slowest chunk we saw was 63.3 seconds, so budget closer to 12 or 13 minutes.