How to

How to transcribe a research interview: consent, pseudonyms and coding

How to transcribe a research interview without breaking what you promised: consent wording, pseudonyms, and a file your coding software can actually read.

To transcribe a research interview, start with the document you handed your participant, not with the audio file. The participant information sheet and the consent form say who may hear the recording, and that one sentence usually settles whether you may put the file through a third party at all. The rest is ordinary work: produce the text, replace identifying detail with pseudonyms as you go, keep the key somewhere the transcript is not, and give your analysis software a plain file it can read.

Somewhere in the paperwork is a sentence about the recording. It might say the audio will be heard only by the research team, or that an approved transcriber will hear it under a confidentiality agreement. Whatever yours says, that is the promise, and it was made to a person rather than to a regulator.

Disclosing it is part of consent. The World Medical Association’s Declaration of Helsinki, revised in 2024 and written for medical research, lists the provisions protecting privacy and confidentiality among the things a participant must be told in plain language, and states the duty behind it flatly: “Every precaution must be taken to protect the privacy of research participants and the confidentiality of their personal information.”

If the sheet named the research team and nobody else, dropping the file into a hosted service adds a party your participant was never told about. The objection is not how long a company keeps the file. It is that the company was not in the sentence. Under the GDPR that company is generally a processor acting on your instructions, with paperwork attached, as our post on GDPR and meeting recordings sets out.

Ethics approval and your data management plan are the next documents to read

Approval was granted for a described procedure, and where the audio goes is part of that description. Whether a change of route has to go back to your committee is their call, not ours. The one answer that is never right is deciding quietly and writing it up later.

Your data management plan is where the practical answers already live: which drive the audio sits on, who has access, how long each file is kept, what gets deposited at the end. Read it before you start.

The GDPR has research-specific wording worth knowing here, because retention is where it bites. Storage limitation, in Article 5(1)(e), lets data be kept longer where it is processed solely for “scientific or historical research purposes or statistical purposes in accordance with Article 89(1)”. Article 89(1) is what that points at: such processing “shall be subject to appropriate safeguards”, which must put technical and organisational measures in place to respect data minimisation, and may include pseudonymisation. That is permission to keep research data longer in exchange for handling it more carefully, not permission to keep raw audio on a laptop forever.

Pseudonymise inside the transcript, and keep the key somewhere else

Do the replacing while you edit, not as a pass afterwards. A later pass means a clear-text transcript existed for a while, and you will miss the third mention of the town.

The GDPR defines pseudonymisation as processing personal data so that it “can no longer be attributed to a specific data subject without the use of additional information”, provided that information is “kept separately” and protected against re-identification. The separateness is the mechanism. A spreadsheet named participants.xlsx sitting in the same folder as the transcripts gives you the paperwork of pseudonymisation and none of the protection.

P04 is P04 in every file, every quote and every filename.

What the participant said What goes in the transcript Why it matters
Their own name The participant code, used everywhere The one substitution people remember to make
A colleague’s or relative’s name [colleague], [her son] Third parties never consented to anything
Their employer [a regional hospital trust] Often identifies the person faster than a name
Their town or team [a town in the north-west] Small settings are identifiable from two details
A rare role, diagnosis or event A generalised description Uniqueness is what identifies, not the label
A date that pins an incident [spring 2025] Dates plus a place reconstruct a person

A voice cannot be pseudonymised, so the audio stays identifiable whatever you do to the text, and it belongs in the most protected place you have.

Match the level of detail to the analytic method

Decide what kind of transcript your method needs before you touch the file. The level of detail is a methodological choice, not a formatting preference.

Thematic analysis usually works from readable text, so fillers and stumbles are noise. Conversation analysis sits at the other end: how something was said is the data, so the transcript has to record delivery as well as words, and there are established notation systems for that. We are not reproducing one from memory; take the conventions from the source your method’s literature points you at.

Automatic transcription lands near clean verbatim on its own, so detailed notation is work you add by ear afterwards rather than work a tool saves you. The transcript types and what each costs are in our guide to transcribing an interview end to end.

Getting the text into coding software means a plain file with stable identifiers

Coding packages want a document, not a subtitle file. NVivo 12’s own help page for importing transcripts says they “must be plain text (.txt) in comma- or tab-separated variable format (CSV or TSV) or rich text or Word documents”, and that a table or delimited import must map one field to Timespan and another to Content. Check the current documentation for whichever package your department licenses.

An SRT or VTT export is the wrong shape here. Those files break text into short numbered cues, so a sentence arrives split across two blocks with a stray number in front of it; subtitle and transcript formats sets out the difference.

What to hand the software instead:

  1. One file per interview, named with the participant code and the date, like P04_2026-03-12.txt.
  2. A short header: participant code, date, length, interviewer initials, how it was transcribed and when.
  3. One paragraph per turn, with a short speaker label you typed yourself, such as “I:” and “P04:”, identical in every file.
  4. Timestamps in one consistent format, and straight quotation marks rather than smart ones.

Transcribing a research interview on your own machine

If your consent wording rules out a third party, the route that fits is transcription that never leaves your computer. FreeTranscribe runs OpenAI’s open-source Whisper model in a browser tab on your own graphics card through WebGPU. The file is read from your disk, nothing is uploaded, and with no account there is no processor agreement to chase through the research office.

The limits matter more here than in most uses:

  • There are no speaker labels. The interview comes back as one undivided block, which is a real problem for a method built on speaker-attributed text. You type the labels in yourself, as described in what speech recognition does and does not do with two speakers.
  • Desktop Chrome or Edge with a working WebGPU adapter. Firefox, Safari and phones are not supported yet.
  • English only for now, so multilingual fieldwork is out of scope.
  • It runs the base English model, about 200 MB downloaded once and then cached by the browser.
  • About 1.5x real time on a desktop with a graphics card in our test, closer to the length of the recording on a thin laptop.
  • Names, place names and technical vocabulary are the weak spots, along with heavy accents and noisy rooms. Read the text against the audio before you code from it.

Keeping the audio local shrinks the list of parties who touch it and changes nothing else. A transcript on your laptop is exactly as much personal data as one on a server, and interviews with no cloud provider in the chain counts the copies.

To try it on a recording you already have, open the home page and drop the file in.

Frequently asked questions

Can I use an online transcription service for research interviews? That depends on what your participants were told and what your approval says, rather than on the service. Where a protocol allows an approved transcriber under a confidentiality agreement, that decision is already recorded, and you follow it.

Should I pseudonymise the recording as well as the transcript? You cannot, because a voice identifies a person. Treat the audio as the identifiable copy: strongest storage, shortest retention, gone once the transcript is checked unless a deposit agreement says otherwise.

Does automatic transcription give me a verbatim transcript? Not in the methodological sense. What comes back carries no notation of pauses, overlap or delivery, so if your method needs that detail, budget for a listening pass you do yourself.

Where do I keep the code that links P04 to a real name? Somewhere the transcripts are not, with its own access control, and named in your data management plan. Article 4(5) hangs pseudonymisation on that information being kept separately.

This is general information, not legal advice. Your own ethics approval, your institution’s data protection policy and whatever you promised your participants govern what you may actually do.

Sources, checked 15 September 2026

research interviewsqualitative researchpseudonymisationconsentlocal transcription
FreeTranscribe

Written by the people who build FreeTranscribe. We test every claim on our own files and date every price. About the site.

Transcribe a file now. Free, in your browser.
Open the transcriber