Does microphone quality or microphone distance matter more?
The signal-to-noise ratio degrades with distance. A laptop microphone twenty inches from a speaker in a quiet room beats a £200 microphone six feet away in an office, because the closer recording has proportionally less background noise in it.
A lavalier or headset microphone clipped to the speaker is the single most effective upgrade available, and it costs very little. It removes the distance problem entirely rather than trying to filter it afterwards.
One microphone in the middle of a table beats several spread around a room. Multiple distant microphones each capture the room as much as the person, and overlapping audio is much harder to transcribe than a single clean source.
How do I reduce background noise in a recording?
Steady background noise is the most damaging case, because the model treats it as part of the speech rather than as something to ignore. Traffic, air conditioning, and a room hum all do this.
Software noise reduction helps more than people expect, but it works by guessing what the speech should sound like, and it introduces artefacts that the model then reads as garbled words. Mild hum benefits; heavy noise often does not.
Turning off air conditioning or moving away from a road takes ten seconds and is more effective than any amount of processing afterwards. Record somewhere quiet rather than recording a quiet recording.
What happens when two people talk at once?
When two people talk at once, the model has to separate interleaved speech, and it does not have enough information to do it reliably. The result is not clean speaker labels but fragments and mis-attributed lines.
The only reliable fix is to prevent it. A conversation with one microphone should be moderated so people speak in turns, which is worth doing anyway in any recording meant to be used later.
For genuinely unavoidable overlap, transcribing the same passage twice does not help, because the model sees the same problem both times. Splitting speakers onto separate tracks and transcribing each separately does.
Does MP3 file format affect transcription accuracy?
Lossy compression discards detail unevenly across the frequency range, and what it discards is concentrated where quiet speech and consonants live. That is precisely where transcription is hardest, so the loss lands where it costs most.
Record to WAV where you can, or at least 192 kbps. Below 64 kbps, typical of phone voice memos, expect errors even on a well-spoken recording.
Mono is usually better than stereo for voice, since stereo doubles the data without adding information and introduces phase problems when the channels are not perfectly aligned. The MP3 to text guide covers the encoding detail.
What should I check before uploading a recording?
Trim leading and trailing silence. It adds no text and occasionally costs you the first words of a file.
Cut any section with unusable audio rather than hoping it improves. A five-second gap of static transcribes as confident nonsense, and those invented words are harder to spot than an obvious omission.
Speak names and technical terms clearly at least once in the recording. The model cannot infer a spelling it has not heard, and this is the most common source of checkable errors in otherwise good transcripts.
Read the transcript against the audio once before publishing anything externally. Names, numbers, and titles deserve the closest attention, because a wrong figure in a correct method is worse than a garbled word.