How to transcribe research interviews without uploading them
Interview transcription for research: verbatim or intelligent verbatim, speaker labels, anonymising as you go, and getting the file into your analysis tool.
A research interview is not a voice memo. It has two or more voices, a consent form behind it, an ethics approval that says where the data may live, and an analysis step waiting downstream that expects a particular shape of file. This guide covers the decisions that actually change the transcript, the mechanics of producing it, and how to keep participant audio off the internet while you do it.
Decide the transcription style before you start
Three styles are in common use, and switching halfway through a study makes your data inconsistent.
- Verbatim keeps every word, including false starts, repetitions, filler words and marked pauses. Choose it when how something is said carries meaning: conversation analysis, discourse analysis, anything where hesitation is data.
- Intelligent verbatim keeps every substantive word and drops the noise: "um", stutters, the third repetition of a word. This is the common default for thematic and content analysis, and it is far easier to read.
- Edited or clean tidies grammar and reorders clauses for readability. It is fine for quotes in a report and wrong for analysis, because you can no longer tell what the participant actually said.
Write the choice into your protocol so a second coder produces the same thing you do.
Set up the recording so the transcript is possible
Transcription quality is decided while you record, not afterwards.
- One room, one recorder, close to the speakers. A phone lying between two people beats a laptop microphone across a table.
- Say the room aloud. Start the recording with the date, the study code and the participant pseudonym. It becomes the first line of the transcript and identifies the file forever after.
- Ask each person to say a sentence at the start. A few seconds of each voice on its own makes speaker separation much more reliable.
- Avoid crosstalk where you can. Any system, human or software, struggles when two people speak over each other.
- Record uncompressed if you have the space. Heavy compression removes exactly the detail speech recognition needs.
Produce the first pass
Typing an interview by hand runs four to six hours per recorded hour. Almost nobody does that any more; the practical method is a machine first pass followed by a human correction pass.
What the first pass has to give you, or you will pay for it later:
- Speaker separation. The transcript should already be split by voice, so correcting it means renaming "Speaker 1" to "P07", not reading the audio again to find the turns.
- Timestamps you can click. Every claim you later quote needs to be checkable against the audio in one action.
- A plain export. Your analysis tool, not your transcription tool, is where the coding happens.
Doing that on the device rather than through a web service matters here for a reason that has nothing to do with preference: participant audio is personal data, and most ethics approvals and data management plans either forbid uploading it or require a data processing agreement with whoever receives it. A tool that never sends the audio anywhere removes the question.
Correct the pass, in one sitting per interview
Work with the audio playing and the text in front of you.
- Fix names, places, technical terms and acronyms first. They are what search and coding depend on.
- Apply your chosen style consistently: if you are doing intelligent verbatim, remove fillers everywhere, not just where they annoy you.
- Mark what the words cannot carry: [laughs], [long pause], [inaudible 00:14:22]. Inaudible markers with a timestamp are more honest than a guess.
- Do not silently correct grammar in a verbatim transcript. Participants do not speak in sentences and that is not an error.
Anonymise as you go
Replace direct identifiers in the transcript at correction time, not in a later pass you will not have time for: names, employers, place names, distinctive job titles, dates of specific events. Keep the mapping between pseudonym and real identity in a separate file, stored separately from the transcripts, and delete the mapping when your protocol says to.
If a participant names someone else, that third party has rights too. Pseudonymise them as well.
Get the transcript into your analysis tool
Most qualitative software imports plain text or Word. What matters is that the export keeps speaker labels as a consistent prefix, one turn per paragraph, so the tool can split turns automatically.
A practical shape that imports cleanly nearly everywhere:
I: What made you change the way you did that?
P07: [00:12:41] It was after the second review. Nobody said it directly, but the message was clear enough.
Consistent prefixes, one blank line between turns, timestamps in a fixed format.
Where ThinkScribe fits
ThinkScribe was built for exactly this shape of work: it records or imports the interview, transcribes it on your own Mac, iPhone or iPad with no server involved, separates the voices, and lets you name a speaker once so that the same person is recognised in every later interview of the study. Every line stays linked to its point in the audio, so checking a quote is one tap, and the transcript exports as TXT, Markdown, DOCX, PDF, JSON, SRT or VTT for whatever comes next. You can also ask questions about a transcript and get answers with citations back to the lines they came from, which is useful for orientation, though it is not a substitute for coding.
Frequently asked questions
What is the difference between verbatim and intelligent verbatim transcription?
Verbatim keeps everything that was said, including fillers, repetitions and marked pauses, and is used when delivery itself is part of the analysis. Intelligent verbatim keeps all the substantive content and removes the noise, which is the usual choice for thematic and content analysis because it is much easier to read and code.
How long does it take to transcribe a one-hour interview?
Typing from scratch takes four to six hours for one hour of audio. A machine first pass followed by a correction pass usually lands between one and two hours per recorded hour, depending on audio quality, accents and how much crosstalk there is.
Do I have to anonymise the transcript?
That depends on your ethics approval and your data management plan, but in practice pseudonymising as you correct is the only version that gets done. Replace names, employers and identifying places in the text, keep the mapping in a separate file, and remember that third parties the participant names have rights too.
Can transcription software label who is speaking?
Yes. Speaker diarization splits the audio by voice and labels the turns, and speaker recognition goes further by matching a voice to someone you have already named. Read How speaker diarization and speaker recognition work for what each one can and cannot do.
Is it safe to upload interview recordings to an online transcription service?
Legally it is a processing decision, not a technical one: the audio is personal data, the service becomes a processor, and you need a lawful basis, a data processing agreement and consent language that covers it. Transcribing on your own device avoids the transfer entirely, which is why on-device tools are the simpler answer for anything covered by an ethics approval.