Accuracy

Improve transcription accuracy: nine ways to record audio that transcribes well

Nine recording habits that improve transcription accuracy: mic distance, one voice at a time, echo, levels, phone placement, file format and a 10-second test.

The quickest way to improve transcription accuracy is to fix the recording, not the transcript. A speech model can only write down what it can hear, so a microphone within arm’s reach, one voice at a time, a room without echo, a healthy level and the original file do more for you than any setting in the tool. Here are nine habits, each with the reason it matters to a speech model and the practical fix. They apply to any transcription tool, ours included.

1. Get the microphone within arm’s reach

A microphone picks up whatever is loudest where it sits. Move it away from the speaker and the voice gets quieter while the fridge and the traffic stay the same, so the model gets a thinner voice in more noise.

Keep the mic within about an arm’s length of whoever is talking. On a laptop, sit up close, because the built-in mic sits around the keyboard or hinge. On a phone, put it on the table between the people talking, never in a pocket. A cheap wired headset or lapel mic beats an expensive mic across the room, simply because it stays close.

2. Let one person speak at a time

The model writes a single stream of words. When two people overlap, the two voices blur into one signal that matches neither, and the model has to pick one or skip the moment. The model’s own paper doesn’t mention overlapping speech, so don’t expect it to cope.

Agree at the start of a meeting that people finish before the next person starts. In an interview, wait a beat after the answer before the next question. If two people do talk over each other, have one repeat the point once the other has finished. That repeated sentence is worth more than any editing later.

3. Take the echo out of the room

Hard walls, glass and bare floors reflect the voice back a few milliseconds late, so the microphone hears every consonant twice, slightly smeared. Consonants are what separate similar words, and echo is what blurs them.

Record in the smallest room with the most soft stuff in it. Curtains, a rug, a sofa and a bookshelf all absorb reflections. Move the mic closer and the direct voice wins over the reflections. A parked car is a surprisingly good booth. A meeting room with a glass wall and a bare table is one of the worst.

4. Turn off the background noise and the music

The model behind FreeTranscribe, OpenAI’s open-source Whisper model, was trained on 680,000 hours of audio that its paper describes as “a very diverse dataset covering a broad distribution of audio from many different environments, recording setups, speakers, and languages”. That’s why it copes with some noise. The same paper tested it with white noise and pub noise added, and reports that “all models quickly degrade as the noise becomes more intensive”. Robust to some noise is not immune to it.

Close the window, switch off the fan or the air conditioning for the length of the recording, and don’t boil the kettle. Music is worse than noise, because words in the music bed are speech to the model, and it will happily transcribe the chorus. Add music in the edit. Our post on invented words in silence explains what a music bed under a pause does to the output.

5. Set the level so it never clips and never whispers

Two opposite mistakes, same result. Too high, and the loudest moments are flattened into distortion, so a consonant no longer sounds like the one the model learned. Too low, and speech sits close to the noise floor, so a quiet sentence and a quiet pause look alike and the model guesses at both.

Watch the meter while someone talks at normal volume. The peaks should sit in the upper part of the meter without touching the top or turning red. Then have them laugh or raise their voice, because that is the moment that clips. If the meter barely moves, move the mic closer rather than turning up the gain, since gain raises the noise along with the voice.

6. Put the phone on the table and into airplane mode

A phone in a hand or a pocket records every rustle and tap louder than the voice, because the mic is touching the thing making the noise. A phone on a bare table picks up every knock through the wood. And a phone still on the network can buzz halfway through the take.

Put the phone face up on something soft, a folded jumper or a notebook, with the microphone end toward the people talking. Switch on airplane mode and do not disturb for the length of the recording. Check that the battery will last and that locking the screen doesn’t stop the recording app, which is what the test in habit 9 is for.

7. Keep the original file and don’t re-compress it

Every pass through a lossy format such as MP3, M4A or AAC throws away a little of the audio, and quiet high-frequency detail, which is where consonants live, tends to go first. Converting a file to make it smaller, or sending it through a messaging app that re-encodes it, hands the model a worse copy of the same speech. A bigger file doesn’t help either: the paper says “all audio is re-sampled to 16,000 Hz” before the model hears it. The goal is the untouched file, not a bigger one.

Upload the file your recorder produced. FreeTranscribe reads MP3, WAV, M4A, AAC, FLAC, OGG and video containers (MP4, MOV, WEBM, MKV and AVI with MP3, PCM or AAC audio). It only reads the audio track, so you don’t need to strip the video out first, and there’s no length limit, so there’s no reason to shrink a long recording.

8. Say who’s speaking and pause when the speaker changes

The model writes words, not names. It doesn’t label speakers, so a three-person meeting comes out as one block of text, and the only way to tell who said what is to have it in the words.

At the start, have each person say their name once in their own voice. When the speaker changes, a pause of about a second gives the model a clean boundary instead of a run-on, and if you export SRT or VTT the cue breaks tend to fall at the pause. The chair saying “over to Sam” is a speaker label you can search for. Our post on SRT, VTT and plain text explains how to use the timestamps.

9. A 10-second test is the fastest way to improve transcription accuracy

Every problem above shows up in the first ten seconds of a recording. Ten seconds costs nothing to check. An hour recorded with the fan on costs an hour.

Record ten seconds of normal talking in the real position, play it back on headphones and match what you hear against the table below. For a harder check, run the clip through the tool. In our test a 71-second clip came back in 41 to 50 seconds on a desktop with a graphics card, including the model load.

What you hear in the recording What the transcript does Fix
The voice sounds far away, the room sounds close Words dropped, short words guessed Mic within arm’s reach (habit 1)
Two voices at once One speaker vanishes or the two get merged One at a time, repeat the lost sentence (habit 2)
A ringing, hollow tone on every word Similar words swapped, endings lost Smaller room, soft furnishings, mic closer (habit 3)
A steady hum, hiss or music bed Extra words, especially in pauses Fan off, window shut, music added later (habit 4)
Crackling on loud words Loud phrases mangled or missing Lower the level, test with a laugh (habit 5)
Very faint voice with hiss behind it Quiet sentences skipped or invented Mic closer, not gain up (habit 5)
Taps, rustles, a buzz halfway through Gaps and stray fragments Phone on something soft, airplane mode (habit 6)

A word on what this tool can and can’t do. FreeTranscribe runs the base English version of OpenAI’s open-source Whisper model on your graphics card through WebGPU, so it works in desktop Chrome or Edge today, not yet in Firefox, Safari or on a phone, and English only. Names, technical terms, heavy accents and noisy rooms are the base model’s known weak spots. Nothing is uploaded, there’s no account and no length cap. If you have just recorded your 10-second test, drop it on the home page and read what comes back before you record the real thing.

Frequently asked questions

Does a better microphone matter more than a better model? For most recordings, yes. A larger model is better at telling apart similar-sounding words in clean audio, but no model can separate two overlapping voices or restore a consonant that was clipped. Fix the recording first.

Can I rescue a bad recording in an audio editor afterwards? Partly. Normalising a quiet file, trimming silence and cutting a music intro all help. You can’t un-mix two people who spoke at once, and you can’t undo clipping, because the flattened part of the waveform is gone. Noise reduction tends to remove speech detail along with the noise, so try the file untouched first.

Does recording in WAV give better results than MP3? WAV avoids one lossy pass, which is a small gain. The bigger gains are mic distance and the room. Whatever you record, the model resamples it to 16,000 Hz, so a 96 kHz file is not heard in any more detail than a 16 kHz one.

Will the transcript say who spoke? Not on its own. FreeTranscribe exports TXT, SRT and VTT, none of which carry speaker names. People saying their names and pausing at each change of speaker (habit 8) is how you get that into the text.

Sources, checked 15 September 2026

  • https://arxiv.org/html/2212.04356 - Radford et al., “Robust Speech Recognition via Large-Scale Weak Supervision”, HTML version: section 2.1 (680,000 hours, “very diverse dataset” quote, 30-second segments), 2.2 (audio re-sampled to 16,000 Hz), 3.7 (white noise and pub noise test, “all models quickly degrade as the noise becomes more intensive”), and the absence of any treatment of overlapping speech
FreeTranscribe

Written by the people who build FreeTranscribe. We test every claim on our own files and date every price. About the site.

Transcribe a file now. Free, in your browser.
Open the transcriber