
Table of contents
- The five levers, ranked by measured effect
- Lever 1: microphone distance — the fivefold lever
- Lever 2: overlapping speakers
- Lever 3: noise, music, and silence
- Lever 4: language, accent, and vocabulary
- Lever 5: the model variant — the only pure software lever
- What doesn’t help much
- Measure the fix, not the feeling
- Frequently asked questions
- Final thoughts on improving transcription accuracy
- Sources
Quick answer: Transcription accuracy is mostly decided before any software runs. The published evidence ranks the levers in this order: microphone distance and room audio (the same flagship model that scores 2–3% word error rate on close-miked speech scores around 16% on distant-microphone meetings — a fivefold penalty), overlapping speakers (the main reason meeting corpora sit at the bottom of every benchmark), noise and long silences (silence is where speech models hallucinate text nobody said), language and vocabulary (published per-language error rates span more than a factor of ten), and finally the model variant — the one lever that is purely a software choice. Everything below links each claim to its measured number.
This guide was written by the LoroNote team on August 27, 2026. Every benchmark figure in it comes from the primary sources collected in our verified transcription statistics page; the measurement method at the end comes from our word error rate guide. LoroNote runs Whisper Large V3 Turbo on the device, so improving your recording conditions improves our own product’s output — we have no incentive to tell you accuracy is fine as it is.
The five levers, ranked by measured effect
| Lever | The measured evidence | What you control |
|---|---|---|
| 1. Microphone distance | ~2–3% WER close-miked vs ~16% on distant-mic meetings (same model) | Where the phone sits |
| 2. Overlapping speakers | Meeting corpora with cross-talk anchor the bottom of every benchmark | Turn-taking, seating |
| 3. Noise and silence | Models degrade sharply below ~10 dB SNR; hallucinations concentrate on silence | Environment, trimming |
| 4. Language and vocabulary | Published per-language WERs span 4.3% to 55.7%; names and jargon fail first | Custom dictionary, clear names |
| 5. Model variant | Whisper Small scored roughly double the errors of Large-class engines in one run | Which app — the only pure software lever |
The order matters. People shopping for a better app are usually reaching for lever 5 while levers 1–3 do the damage.

A dynamic stage microphone — but the lever is not the hardware: it is how close any microphone, including the one in your phone, gets to the voice. Source:
Wikimedia Commons (E bailey)
, CC BY-SA 4.0.
Lever 1: microphone distance — the fivefold lever
The single largest accuracy factor is how far the microphone sits from the mouth. The cleanest published demonstration is one model measured across conditions: Whisper Large V3 averages about 7.4% WER across eight varied real-world English test sets, roughly 2–3% on clean close-miked read speech — and around 16% on the AMI corpus, which was recorded in meeting rooms with distant microphones. Nothing about the software changed between those rows.
What to actually do:
- Move the phone, not the budget. A phone lying near the person speaking beats an expensive setup across the room. For a two-person interview, the phone belongs between you, closer to the quieter voice.
- In meetings, centre the device on the table rather than at your seat. Every extra metre to the farthest speaker costs more than any app switch will recover.
- For solo dictation, you are already winning. Arm’s length into a phone is close-miked speech — the top rung. If your solo recordings transcribe badly, your problem is one of the levers below, not distance.
Lever 2: overlapping speakers
Cross-talk is the second reason meeting audio anchors the bottom of the benchmarks. When two people speak at once, there is no clean signal to transcribe — the model must pick one stream, and every word of the other becomes an error. This is a physics problem before it is a software problem, which is why it survives every model upgrade.
What to actually do:
- One voice at a time is worth more than any setting. In meetings you control, the discipline of finishing sentences is an accuracy feature.
- Seat the important voices near the microphone. If a decision-maker mumbles from the far corner, that corner is where your transcript will fail.
- Expect speaker labels to degrade with overlap too. Identifying who said what is a separate step with its own failure modes — how speaker identification works covers where it breaks. A transcript can get the words right and still attribute them wrongly during cross-talk.
Lever 3: noise, music, and silence
Steady background noise degrades transcription gradually; the Whisper paper’s own robustness tests show recognition collapsing once additive noise passes roughly the 10 dB signal-to-noise mark — the level of a loud café or bar, where competing models degraded even faster than Whisper. Music is worse than noise for a subtler reason: speech models trained on huge web corpora have seen a lot of music with lyrics, and they will try to transcribe it.
Silence carries its own trap. A 2024 study of Whisper’s API found hallucinated phrases in roughly 1% of transcriptions — invented sentences concentrated where the audio goes quiet, and disproportionately in the speech of people who pause longer. (Silence-born insertions are also how a word error rate can exceed 100%.)
What to actually do:
- Choose the quieter room over the better microphone. A phone in a quiet room beats a good mic in a loud one.
- Turn background music off when a recording matters. It is both masking noise and competing text.
- Trim long silent stretches before transcribing, or record in segments — silence is not neutral padding; it is where invented text appears. Then read the quiet moments of the transcript with suspicion.
Lever 4: language, accent, and vocabulary
Published per-language figures for the same Whisper model span more than a factor of ten — from 4.3% (Dutch) to 55.7% (Albanian) on Common Voice 15. Accent gaps are real and documented too: peer-reviewed research measured an average 35% error rate for Black American speakers versus 19% for white speakers across five commercial systems in 2019-era tests. You cannot configure your way out of a model’s training distribution — but two things are genuinely in your control:
- Names, jargon, and product terms fail first and repeat forever. A model has never seen your client’s name; it will guess the same wrong word every time. This is the one accuracy problem with a direct fix: LoroNote’s custom dictionary exists to teach recurring terms once, and the model stops guessing.
- Say numbers, names, and codes deliberately. The habit costs nothing and targets exactly the words whose errors cost the most — a mangled date and a mangled filler word count the same in a benchmark, but not in your day.
- If your language sits low on the published charts, a two-minute test of your own tells you more than any table — the per-language chart and why some languages are scored by CER instead of WER are in the WER guide.
Lever 5: the model variant — the only pure software lever
Once the audio is as good as life allows, the remaining lever is which model your app runs. The spread is not subtle: in the one benchmark that ran them side by side, Whisper Small scored 3.74% on LibriSpeech test-clean and 7.95% on test-other, while Large-class engines sat near 2% and 4.5% — roughly half the errors, from a model-size choice the app made for you. “Powered by Whisper” on an app listing covers sizes whose accuracy differs by a factor of two, and several iOS apps quietly run the small sizes.
What to actually do: before switching apps over accuracy, check which variant each app actually loads. LoroNote runs Whisper Large V3 Turbo — a Large-family model — on the device; on an iPhone 15 Pro that transcribes at roughly 19× real time (our own measurement), so the large-model accuracy does not come with a coffee-break wait.
What doesn’t help much
Three fixes people reach for that the evidence does not support:
- Recording at a higher sample rate. Whisper resamples all audio to 16,000 Hz before it listens — the paper says so explicitly. A 96 kHz studio file and a plain voice memo arrive at the model looking the same; spend the effort on distance and quiet instead.
- A newer phone. The microphone in any recent iPhone is not the bottleneck; its position is. (A newer phone can transcribe faster, which is a different property.)
- “Enhancement” filters before transcription. Aggressive noise-removal can smear the harmonic structure the model reads. If you use one, verify it with a before-and-after count on the same clip rather than trusting the cleaner-sounding audio — cleaner to your ear is not cleaner to a spectrogram reader.
Measure the fix, not the feeling
Every lever above is testable in ten minutes: record two minutes in your current setup, transcribe, count errors in a 200-word stretch with fixed rules; change one thing — move the phone, kill the music — and run the same passage again. The full counting method, including deciding up front what counts as an error, is in our WER guide. Because LoroNote transcribes on the device and the first two minutes of every recording are free, the before-and-after experiment costs nothing and works with Airplane Mode on.
Frequently asked questions
Why is my transcription so inaccurate?
Check the levers in order: how far the microphone was from the speakers, whether people talked over each other, how loud the room was and how much silence the recording contains, whether the failing words are names and jargon, and finally which model your app runs. In published benchmarks, the distance from close-miked to distant-microphone audio alone moves the same model from roughly 2–3% to around 16% word error rate — more than any app switch recovers.
Does an external microphone improve transcription accuracy?
Distance improves accuracy; an external microphone is one way to buy distance. A lapel mic on the speaker is effectively close-miked speech even in a big room. But a phone placed close achieves most of the same effect for free — position first, hardware second.
Does background music affect transcription?
Yes, twice over: it masks the voice like any noise, and because speech models have seen huge amounts of lyric-bearing audio, they may transcribe the song instead of the speaker. Recognition also degrades sharply once noise passes roughly the 10 dB signal-to-noise level of a loud café. Turn music off for recordings that matter.
Can I fix accent errors with a setting?
Not with a setting — accent performance lives in the model’s training data, and research has documented real demographic gaps. What is in your control: choose an app running a Large-class model (bigger models absorb more variation), teach recurring names and terms to a custom dictionary, and measure on your own two minutes rather than trusting averages measured on other people’s voices.
Do longer recordings transcribe less accurately?
Length itself is not a documented accuracy factor — condition is. A three-hour lecture recorded close and quiet transcribes like its first ten minutes. What grows with length is the cost of reviewing errors, which is why the fixes above matter more the longer your recordings run.
Final thoughts on improving transcription accuracy
Accuracy advice usually points at software because software is easy to switch. The measured evidence points the other way: the fivefold penalty lives in microphone distance and room audio, the stubborn failures live in cross-talk and silence, and the one honest software lever is the model variant your app loads. Move the phone, quiet the room, take turns, teach it your names, check the model — in that order — and verify each change with a counted before-and-after on your own recording, because the only accuracy number that matters is yours.
Sources
- AI transcription statistics: every number verified — the primary sources and verification dates for every benchmark figure used above
- OpenAI, Robust Speech Recognition via Large-Scale Weak Supervision — 16,000 Hz resampling; additive-noise robustness (degradation below ~10 dB SNR pub noise)
- Open ASR Leaderboard — mixed-benchmark and AMI meeting-room figures
- Koenecke et al., Racial disparities in automated speech recognition (PNAS, 2020) — the measured accent/dialect gap
- Koenecke et al., Careless Whisper: Speech-to-Text Hallucination Harms (2024) — hallucination on silence
- Lyonesse, Apple’s New Speech API vs Whisper — the side-by-side Small vs Large-class run
- LoroNote, What is word error rate? — the measurement method for before-and-after tests
Figures re-checked against the linked sources on August 26–27, 2026.
Turn your voice into text — offline
LoroNote transcribes meetings, lectures, and interviews right on your iPhone — private, accurate, and unlimited.
Download on the App Store