Guide

Turn a podcast episode into a transcript, show notes and chapters

A repeatable workflow for podcast transcription: clean speaker-labelled text, show notes and chapter markers, subtitles for the video cut, all from one recording.

Podcast transcription is one recording feeding five outputs: the transcript on your episode page, the show notes, chapter markers, subtitles for the video cut, and the pull quotes you post. Doing all of that by hand takes longer than recording the episode. This guide sets out a workflow that produces every one of them from a single pass, and says where each output actually goes.

Why publish a transcript at all

Three reasons, in the order they pay off.

  • People find you through words, not audio. Search engines index the transcript on your episode page; they cannot index an MP3. A guest's name, a product they mentioned, an argument you made forty minutes in: all of it becomes findable.
  • Some of your audience needs it. Deaf and hard of hearing listeners, people in a noisy commute, non-native speakers who read faster than they listen, and anyone who wants to skim before committing an hour.
  • It is the source for everything else. Show notes, chapters, quotes and social posts are all edits of the transcript. Produce it first and the rest is selection, not transcription.

Start from the right audio

Feed the transcription the cleanest audio you have, which is usually not the final published mix.

  • Use the edited master, not the raw multitrack. The edit already removed the retakes and the dead air.
  • If you record each speaker on a separate track, keep them. Per-speaker tracks make speaker labels exact rather than inferred, and they survive crosstalk.
  • Skip the music bed if you can. Intro music over speech is the single most common cause of garbled opening lines. Transcribe the speech-only version and paste the intro copy in by hand.
  • Export at a sane quality. A heavily compressed MP3 loses the detail speech recognition uses. WAV or a high-bitrate M4A is better input.

Produce the transcript

Import the file and let it transcribe. What you want out of this step:

  1. Speaker-labelled turns. Two or three named voices, not one wall of text. If your tool recognises voices between episodes, name your co-host once and every future episode arrives already labelled.
  2. Timestamps on every line, so a line and its moment in the audio stay connected. Chapters and clip selection both depend on this.
  3. Correct names. Guest names, company names and jargon are exactly what people search for and exactly what speech models get wrong. Add them to a custom vocabulary before the run when your tool supports it, or fix them once with a find and replace afterwards.

Then read it once with the audio playing at 1.5x, fixing names and any sentence that changes meaning. Do not tidy every filler word; a podcast transcript is allowed to sound like people talking.

Cut the show notes from the transcript

Show notes are a summary plus a list of references, and both are already in the transcript.

  • Summary: three to five sentences that say who is on, what the argument is, and why someone should spend an hour. Write it after you have the transcript, not from memory.
  • Links and mentions: scan the transcript for every book, tool, person and paper named. That list is the part listeners actually come back for.
  • Key moments: five to eight lines with timestamps. These double as chapters.

An on-device assistant that can read the whole transcript will draft the summary and pull the mentions for you. Treat its output as a first draft with the timestamps attached, then cut it down. The check that matters is whether each claim is in the transcript.

Chapter markers

Most podcast apps read chapters from the episode file or the show notes, and YouTube builds them from timestamps in the description. The format that works nearly everywhere is a line per chapter, starting at zero:

00:00 Cold open
01:42 How the study was designed
14:05 The result nobody expected
28:30 What it means for practitioners
47:11 Where to follow the work

Take the timestamps from the transcript rather than scrubbing the audio. Round down to the sentence that starts the topic, never up: landing a listener slightly early is fine, landing them mid-sentence is not.

Subtitles for the video cut

If you publish video or clips, export subtitles from the same transcript rather than letting each platform auto-caption. You control the spelling of names, and the same file works on every platform.

  • SRT for YouTube, LinkedIn and most editors.
  • VTT for web players and anything embedded in a page.

Keep lines short, roughly forty characters, two lines at a time, so captions do not cover the frame. How to create SRT subtitles from an audio file on Mac covers the mechanics and the timing fixes.

Put the transcript on the episode page

Publish the text on the page itself, not as a downloadable file, or search engines get nothing from it. Structure it the way it was spoken: speaker name, then the turn, one paragraph each. Timestamps every few minutes help readers jump; a hidden wall of text helps nobody.

Doing it without uploading the episode

Everything above works with cloud services. It also works entirely on your own machine, which matters if the episode is under embargo, the guest spoke off the record before the edit, or you would rather not upload an unreleased file at all.

ThinkScribe runs the whole pass on your Mac: import the master, get a speaker-labelled transcript with word-synced playback, ask for a summary and action points from a model running locally, and export the result as text, Markdown, PDF, DOCX, SRT or VTT. Name a voice once and your co-host is recognised in every later episode.

Frequently asked questions

How do I transcribe a podcast episode?

Export the edited master as WAV or high-bitrate M4A, transcribe it with speaker labels and timestamps, then do one correction pass focused on names and jargon. From that corrected transcript you cut show notes, chapters and subtitles without transcribing anything twice.

How long does it take to transcribe an hour-long episode?

The machine pass is usually a few minutes on a recent Mac. The correction pass is the real cost: budget twenty to forty minutes for an hour of two-person conversation, most of it spent on names, brands and technical terms.

Should podcast transcripts be verbatim?

No. Readers want an intelligent verbatim transcript: everything substantive, minus the false starts and most of the fillers. Keep the character of how people speak, remove the noise that makes the page hard to read.

How do I add chapters to a podcast?

Write one line per chapter as a timestamp and a short title, starting at 00:00, and put that list in the episode description or the show notes. Take the timestamps from your transcript, rounding down to the sentence where the topic starts.

Can I transcribe a podcast without uploading it?

Yes. On-device transcription runs the speech model on your own computer, so an unreleased episode never leaves it. That also removes the per-minute billing, which matters when you are transcribing a weekly show plus outtakes.

Try the private alternative.

Free to download, with free uses of every Pro feature. No account needed.

Download on the App StoreDownload on the App Store

Audio to text →