How to transcribe an interview: the whole job, end to end
How to transcribe an interview end to end: choose verbatim, clean verbatim or notes, record two people well, run the file, then edit and timestamp quotes.
To transcribe an interview, decide first what kind of transcript you need, because that one decision sets how much work every later step takes. Then record the two of you so a machine can hear both voices, run the file through a transcription tool, and spend one focused pass fixing names and marking the times of anything you plan to quote. The recording and the editing are where the real effort sits. The transcription itself is mostly waiting.
Decide which transcript you need before you transcribe an interview
People say “I need this interview transcribed” as though there were one thing to produce. There are at least three.
A full verbatim transcript keeps everything: the ums, the false starts, the abandoned half-sentence, the laugh. You want it when how something was said matters as much as what was said.
Clean verbatim keeps the words as spoken but drops the noise around them. The fillers go, the stutter becomes one clean word, “you know” disappears when it is a tic. Nobody’s sentences get rewritten.
Notes are not a transcript at all. You summarise the conversation in your own words and type out exactly, in quotation marks, only the passages you might use. For an hour of tape you want three quotes from, this is often the honest answer.
Those labels are working shorthand, not a published standard, so if someone commissions a transcript from you, ask which they mean before you start.
| Transcript type | What it keeps | Pick it when | What it costs you |
|---|---|---|---|
| Full verbatim | Every word, filler, stumble and repeat | The delivery matters, or someone may check your text against the tape | The longest edit, because you add back what the model tidied away |
| Clean verbatim | The words as spoken, fillers and stumbles removed | Publishing quotes, internal write-ups, most interviews | A moderate edit, mostly deletion |
| Structured notes | Your summary, plus exact quotes in quotation marks | You know you want a few quotes from a long conversation | The least typing, but you cannot re-use it later as a record |
Automatic transcription lands near clean verbatim on its own, so full verbatim means putting the mess back by ear, which is slower than removing it.
Recording two people so there is something worth transcribing
An interview is the easiest recording situation there is, and people get it wrong in the same two ways: the recorder sits too far from the person answering, and the room is too hard.
Put the microphone closer to your guest than to yourself. This feels rude and it is correct. You know what your questions were, and a mangled question costs you nothing, while a mangled answer costs you the quote. On a phone, that means the phone nearer to them, microphone end pointing their way, on something soft.
Sit at the corner of a table rather than across the width of it. A corner puts two mouths a similar distance from one device and removes the expanse of reflective table between you. Small room, soft furnishings, window shut. Our post on recording audio that transcribes well has the rest.
One more decision at the recorder: if your device can write two tracks, one per microphone, turn that on. Two files means two transcripts, each already belonging to one person, which saves the labelling job below. And leave a beat at each turn change. A deliberate half-second before your next question becomes a visible gap in the timings, and that gap is what tells the voices apart.
Running the file, and what actually comes back
This step is the least interesting part, which is how it should be. Drop the file in, wait, read what comes back.
FreeTranscribe runs in a browser tab: the speech model runs on your graphics card through WebGPU, the file is read straight from your disk, and nothing is uploaded. It is free, there is no account and no length cap, and it exports TXT, SRT and VTT. The other routes are laid out in our overview of ways to convert audio to text.
The limits matter for interview work, so here they are before you commit an afternoon:
- Desktop Chrome or Edge with a working WebGPU adapter. Firefox, Safari and phones are not supported yet. English only for now.
- It runs OpenAI’s open-source Whisper model in the base English size, about 200 MB downloaded once and then cached by the browser.
- About 1.5x real time on a desktop with a graphics card in our test. A thin laptop takes roughly as long as the recording, so a 90-minute interview is a 90-minute job you can walk away from.
- Names, job titles, company names and technical terms are the known weak spots, along with heavy accents and noisy rooms. Read it against the audio before you rely on it.
- There are no speaker labels. This matters a lot for an interview, so be clear-eyed about it: your conversation comes back as one undivided block with nothing marking who spoke. The Whisper paper puts speaker diarization among the components a fully featured system can involve, and says those are “often handled separately, resulting in a relatively complex system around the core speech recognition model”, so no setting turns it on. Adding the labels yourself takes a few minutes if you work from the timings, and what speech recognition does and does not do with two speakers walks through that.
The editing pass, in the order that wastes least time
Do these in order. The order is the point, because fixing the same word forty times by hand is what makes people say transcription takes forever.
- Skim it once without the audio. You are gauging how good it is, not hunting errors. Two minutes.
- Fix names first, with find and replace. Your guest, their company, the products and people they mentioned. The model gets these wrong consistently rather than randomly, so one replace fixes every instance. Do it before you read closely, or you will read past them.
- Fix the technical vocabulary the same way. Same logic, same tool.
- Add the speaker labels. Short and identical every time, so find and replace still works: “Q:” and “A:”, or two sets of initials.
- Apply your chosen transcript type. For clean verbatim that is a deletion pass over fillers. For full verbatim it is a listening pass putting them back.
- Break the text into paragraphs at turn and topic changes. A wall of text is unusable even when every word is right.
- Read the passages you intend to quote against the audio. Only those. Nothing automates this.
Our guide to transcript mistakes to fix before you publish covers the errors worth hunting in step seven.
Timestamps are how you find a quote again three weeks later
Export SRT or VTT rather than TXT for any interview you will come back to. TXT throws the timings away, and the timings are what let you jump to a moment instead of scrubbing an hour of audio for a sentence you half remember.
They are close enough to navigate by. The paper describes predicting time relative to the current audio segment, “quantizing all times to the nearest 20 milliseconds”, so a cue start lands you within a fraction of a second of the words.
The habit worth building: during the editing pass, put the start time in square brackets next to every passage you might quote, like [00:14:22]. When an editor asks where a quote came from, you answer in ten seconds. Which format to pick is in our post on subtitle and transcript formats.
Where this workflow stops being enough
More than two speakers is the first case: hand labelling scales badly, and a room of six with interruptions is a different problem. Heavy overlap is the second. When two people talk at once the model writes one stream of text for that window, and no editing recovers the buried sentence. Non-English audio is out of scope.
If you promised a source confidentiality, transcription is one link in a longer chain of copies, which our post on source interviews with no cloud provider works through.
For a two-person interview in English, though, the workflow above is the whole job. Drop the recording on the home page and export SRT. The first run downloads the model, and after that it is cached.
Frequently asked questions
How long does the editing take after the tool has run? It depends on the transcript type and how clean the recording was, so we would rather not invent a ratio. The order in the editing section is what controls it: find-and-replace before close reading, and close listening only on what you will quote.
Do I need to transcribe the whole interview? Often not. If you want three quotes from an hour, notes plus exact quotes is a legitimate output. Full transcripts earn their keep when the material will be re-used, checked by someone else, or searched.
Will the transcript tell me who said what? No. FreeTranscribe returns text and timings with no speaker attribution, so a two-person interview arrives as one block. You add labels from the gaps between cues, or record each person to a separate track.
Can I quote straight from the transcript? Not without checking. Names, technical terms and anything said in a noisy room are where this kind of model slips, and a misquote is the one error that costs you something.
Sources, checked 15 September 2026
- https://arxiv.org/html/2212.04356 - Radford et al., “Robust Speech Recognition via Large-Scale Weak Supervision”: diarization named as a component “often handled separately, resulting in a relatively complex system around the core speech recognition model”, and timestamps “quantizing all times to the nearest 20 milliseconds”
- https://github.com/openai/whisper - openai/whisper README: the task list for the released models, “multilingual speech recognition, speech translation, spoken language identification, and voice activity detection”
- The verbatim, clean verbatim and notes distinction is working shorthand rather than a published standard, so no source is cited for it. FreeTranscribe’s behaviour, formats and speed come from our own testing.