WAV to text — uncompressed audio, without the upload
WAV is the format that makes cloud transcription painful: an hour of CD-quality stereo runs to roughly 600 MB, and that is 600 MB you have to upload before anything begins. LoroNote reads the .wav where it already sits, transcribes it on your iPhone or iPad with Whisper Large V3 Turbo, and exports TXT or SRT. Nothing uploads, so file size costs you storage rather than time.
- Input
- .wav, any bit depth or sample rate
- Output
- TXT, or SRT with timecodes
- Where it runs
- On your iPhone or iPad
- Upload
- None — the file never moves
- Price
- Free download · $39.99 lifetime
How to transcribe a WAV file
- 1
Install LoroNote and open it
The download is free and no account is created — there is no sign-up screen, because there is no server holding your transcripts.
- 2
Tap Import File
Get the .wav onto the device first if it came off a recorder — AirDrop, a Lightning or USB-C card reader, or iCloud Drive all work — then pick it in the file browser. There is no size threshold that pushes you into a different workflow.
- 3
Confirm the language
LoroNote detects the spoken language automatically and supports over 100. If the recording is full of names, product codes, or jargon, add them to the custom dictionary first so they come back spelled correctly.
- 4
Let it transcribe on the device
Whisper Large V3 Turbo runs on the phone or iPad itself — airplane mode changes nothing. Keep the app open while a long file processes; it uses the device GPU and pausing it costs you time.
- 5
Review the speakers, then export
Check the speaker labels, rename them, then export. Field recordings of interviews are the case where SRT earns its keep: every line carries a timecode you can quote against.
Where WAV files come from
Nobody ends up with a .wav by accident. It is what you get when someone chose to keep the audio uncompressed:
- Field recorders from Zoom, Tascam, and Sony, which record WAV by default
- Audacity, Logic, and other editors, whose default export is WAV
- Interview and podcast kits that record each microphone to its own track
- Dictation hardware used in medical, legal, and research settings
- The master copy you kept before compressing anything for distribution
How long it takes
Processing time follows the length of the recording, not the size of the file — which is exactly why a large WAV behaves so differently here than it does on a cloud service.
| Recording length | Time to transcribe |
|---|---|
| 5 minutes | about 16 seconds |
| 30 minutes | about 1.5 minutes |
| 1 hour | about 3 minutes |
| 3 hours | about 10 minutes |
Measured by LoroNote on an iPhone 15 Pro at roughly 19x real time. Newer devices transcribe faster; older supported models take longer.
What is different about WAV
The size arithmetic
16-bit, 44.1 kHz stereo works out at roughly 10 MB per minute — about 600 MB for an hour. A 24-bit/96 kHz field recording is more than three times larger again. Check free space before importing a full session.
Uncompressed does not mean clean
WAV faithfully preserves whatever the microphone heard, air conditioning included. Fidelity helps a transcript only up to the point where the room stops being the limiting factor.
A multitrack session is not one file
If your recorder wrote a separate .wav per microphone, transcribe the mixdown or the track with the clearest voice. Speaker identification then separates the people inside that file.
Higher bit depth will not rescue a bad recording
24-bit gives an editor headroom; it does not give a speech model more words. Mic placement is worth more than every setting on the recorder combined.
For WAV, the upload is the whole problem
Cloud transcription hides a step that WAV makes impossible to ignore. Before a server can start, a 600 MB file has to cross your connection — minutes on good office wifi, considerably longer on hotel wifi or a phone hotspot, and a failed upload means starting again. Many services also cap the size of a single file, which is precisely the wall a long uncompressed interview hits.
Transcribing on the device removes the transfer entirely. The recording is read from local storage, processed by the device GPU, and the transcript is written next to it. For the material that usually arrives as WAV — depositions, patient dictation, research interviews, unreleased recordings — never putting the file on someone else’s network is also the point.
Common questions
Is there a file-size limit for importing a WAV?
LoroNote sets no plan cap on file size, length, or number of imports. The practical ceiling is your device: the .wav has to fit in free storage, and a multi-hour session needs enough battery to finish processing with the app open.
Do I need to convert WAV to MP3 before transcribing?
No. WAV imports directly, alongside MP3, M4A, MP4, and MOV. Converting to MP3 first would shrink the file, but since nothing is being uploaded there is no transfer to speed up — you would only be discarding audio.
Can I transcribe a WAV recorded on a field recorder rather than the phone?
Yes. Move the file to the iPhone or iPad first — AirDrop from a Mac, a card reader, or iCloud Drive — then use Import File. Nothing about the workflow changes once the file is on the device.
Does a higher sample rate improve accuracy?
Barely. Speech recognition models work from a downsampled representation of the audio, so 96 kHz and 44.1 kHz recordings of the same conversation transcribe almost identically. Distance from the microphone and room reverberation make a far larger difference.
Are speaker labels available on imported WAV files?
Yes, the same as for a recording made in the app. Speaker identification runs on the device after transcription, labels each segment, and the names you assign survive TXT and SRT export.
Related
- MP3 to text Podcasts, dictaphone exports, and webinar audio, transcribed without an upload or a minute cap.
- Video to text MP4 and MOV: the audio track is read straight out of the file, no editor export first.
- Subtitle generator (SRT) Timed .srt files from audio or video, with speaker labels, generated on the device.
Test it on one session file
Copy a single .wav onto the device, import it, and compare the transcript against what your current workflow produces.
Download LoroNote