How to transcribe a podcast episode and turn it into show notes
How to transcribe podcast episodes free in your browser, then turn the transcript into show notes: summary, chapters with timestamps, links and a quote.
The quickest way to transcribe a podcast episode is to drop the final mixed file into a transcription tool that runs on your own computer, wait about as long as the episode, and export a transcript with timestamps. From there, show notes are a reading job rather than a writing job: the summary, the chapters, the links and the pull quote are already in the text. This guide goes from the exported MP3 to a published episode page, using FreeTranscribe for the transcription step because it’s free and never uploads the file.
Transcribe podcast audio on your own computer
Start with the final mix, the same MP3 or WAV you’re about to send to your host, not the raw session. The final file has the edits, the intro and the ad reads in place, so every timestamp in the transcript matches what listeners hear.
Open the tool in desktop Chrome or Edge and pick the file. The first run downloads OpenAI’s open-source Whisper model, about 200 MB, which the browser caches. The speech recognition then runs on your graphics card through WebGPU. Nothing leaves the machine, which matters when the episode has an embargoed guest or an uncleared sponsor read.
On speed: in our test a desktop with a graphics card ran at about 1.5 times real time, so a 60-minute episode takes around 40 minutes. A thin laptop takes about as long as the episode itself. Leave the tab open. There’s no per-minute charge and no cap, so a three-hour special goes through the same way as a ten-minute update.
The limits, stated plainly: desktop Chrome or Edge only, English only for now, and it’s the base model, so names, jargon and cross-talk are its weak spots. Phones, Firefox and Safari aren’t supported yet.
When it finishes, export twice: TXT for the reading copy you’ll edit into show notes, and SRT or VTT for the timestamps. Keep both in the episode folder. SRT vs VTT vs plain text covers the differences, but for this job either subtitle format works.
Read it through and fix names before anything else
Do the read-through once, at the start, in the TXT file. Every later piece copies from this text, so a guest name fixed here is fixed everywhere.
Check the guest’s name and title against their own site or bio. Then company and product names, place names, anything the guest cited, and every number. Read numbers against the audio. “Fifteen” and “fifty” are easy to confuse, and figures are what listeners quote.
Use find and replace for names that recur. If the guest is Sarah Smyth and the model wrote Sarah Smith 30 times, one replace fixes all of them. Then listen to any passage that reads as nonsense. It’s usually cross-talk, a laugh, or a technical term the base model hasn’t met. Type what was said.
Budget 10 to 20 minutes per hour of audio for this step, more for a panel with two or three guests. It’s the slowest part of the workflow and the one that pays off most.
Pull chapter timestamps from the transcript
Open the SRT or VTT export. Each cue is a few seconds of speech with a start and end time. Your chapters come from those cues.
Find the topic changes in the text, not in the audio. Skimming an hour of transcript takes a few minutes; scrubbing an hour of audio takes far longer. Mark each place the conversation turns: the end of the intro, the first real question, each new subject, the listener questions, the outro. Six to twelve chapters is a comfortable number for an hour.
For each mark, find the same sentence in the SRT and take the start time of that cue. Round down to the whole second, so a listener who taps the chapter lands a breath before the sentence rather than halfway through it. If a topic starts with the host’s question, use the question, not the guest’s answer.
Chapter titles should describe the content, not the structure. “Why the first price was too low” beats “Topic 2”.
Write the show notes from the transcript
Now the writing, and most of it is selecting. Four pieces, each from a different part of the transcript.
The summary, about 150 words. Read the first five minutes and the last five. The intro tells you what the host promised; the outro tells you what they felt the episode delivered. Write the summary between those two: who the guest is, what the conversation is about, and one specific thing a listener will take away. Skip the “in this episode we discuss” opener.
The topic list with timestamps. This is the chapter list from the previous step, written out as bullets: timestamp, then a one-line description. If your host or player supports chapter markers, the same list feeds those too.
The links mentioned. Search the transcript for “dot com”, “website”, “book”, “the paper”, “I wrote” and “you can find”. Each hit is a link. Look every one up and check the URL yourself, because a transcript spells names as they sound.
The pull quote. Search for the moment you remember from the recording, or scan for short declarative sentences from the guest. One or two sentences, in the guest’s own words, that make sense on their own. Trim “um” and “you know” but don’t change the meaning. Put the timestamp after it.
Here is where each piece comes from and roughly how long it takes once the transcript exists.
| What to publish | Where it comes from in the transcript | Rough time budget |
|---|---|---|
| Corrected transcript | The TXT export, read through against the audio | 10 to 20 min per hour of audio |
| Chapter list with timestamps | Topic changes in the TXT, start times from the SRT | About 15 min |
| 150-word summary | The first and last five minutes | About 10 min |
| Links mentioned | Searching the text for “dot com”, “book”, “I wrote” | 5 to 10 min, plus checking each URL |
| Pull quote | One guest sentence that stands alone, with its timestamp | About 5 min |
| Full transcript page | The corrected TXT with speaker names and chapter headings | About 10 min |
These budgets are planning figures, not measurements. A tightly produced solo show comes in under them; a three-person panel that wanders does not.
Publish the full transcript as its own page
There are two reasons to publish the whole corrected transcript, not just the notes.
The first is accessibility. Deaf and hard-of-hearing people can’t use the audio at all. Others prefer reading, are somewhere they can’t listen, or want to check one detail without scrubbing through an hour. A transcript page gives all of them a way in.
The second is search. Search engines index text, and an audio file on its own gives them little to work with beyond the title and description. A full transcript puts every topic, name and term the guest mentioned on a page that can be found. Keep expectations modest: a transcript page doesn’t guarantee rankings, but it does give the episode a text version that can be indexed at all, on your own site.
Apple’s support page for podcast creators says Apple automatically generates transcripts after a new episode is published, for shows in English and ten other listed languages, and that listeners on iOS 17.4 or later can read the full text in the Podcasts app, search it, and tap a line to play from there. The same page says that if you’d rather listeners see your own version, you can supply a VTT or SRT file through the RSS transcript tag and set the episode to “Display transcripts I provide”. Your hosting platform has to put that tag in the feed, so check what yours supports.
For the page itself: episode title at the top, the 150-word summary, then the transcript with the speaker’s name in bold at each turn and a timestamped heading at each chapter. Put it on the same site as the episode page and link the two both ways.
To try the workflow on your latest episode, transcribe it in your browser. It’s free, there’s no account and no length limit, and the file stays on your computer.
Frequently asked questions
How long does it take to transcribe a one-hour podcast episode? In our test, about 40 minutes on a desktop with a graphics card and roughly an hour on a thin laptop, with the tab left open. Add 10 to 20 minutes for the read-through.
Do I need speaker labels? For the show notes, no. For the full transcript page, yes, because a two-voice conversation is hard to follow without them. The export doesn’t include them, so add them during the read-through.
Should I publish the raw transcript or the corrected one? The corrected one. A raw transcript with the guest’s name wrong is worse for the guest than no transcript at all. If you’re short on time, publish the notes now and the transcript once it’s been read.
Can I reuse the SRT file for a video version? Yes. If you publish a video cut of the episode, the corrected SRT works as its captions. How to make an SRT file covers the format and the line-length rules.
Sources, checked 14 September 2026
- Apple Podcasts for Creators, “Transcripts on Apple Podcasts”: https://podcasters.apple.com/support/5316-transcripts-on-apple-podcasts