Offline subtitle generator: audio or video to SRT on your iPhone
LoroNote turns an audio or video file into a timed .srt subtitle file entirely on your iPhone or iPad. Whisper Large V3 Turbo does the recognition locally in over 100 languages, speaker labels survive the export, and the SRT drops straight into YouTube, Final Cut Pro, CapCut, or a media player as a sidecar. Nothing is uploaded and nothing is metered.
- Input
- MP4, MOV, MP3, WAV, M4A
- Output
- .srt with timecodes
- Languages
- 100+, auto-detected
- Internet
- Not required
- Price
- Free download · $39.99 lifetime
How to generate an SRT file
- 1
Install LoroNote and open it
The download is free and no account is created — there is no sign-up screen, because there is no server holding your transcripts.
- 2
Tap Import File
Choose the video or audio file you are captioning. Video works directly — LoroNote reads the audio track — so there is no need to bounce an audio-only version out of your editor first.
- 3
Confirm the language
LoroNote detects the spoken language automatically and supports over 100. If the recording is full of names, product codes, or jargon, add them to the custom dictionary first so they come back spelled correctly.
- 4
Let it transcribe on the device
Whisper Large V3 Turbo runs on the phone or iPad itself — airplane mode changes nothing. Keep the app open while a long file processes; it uses the device GPU and pausing it costs you time.
- 5
Review the speakers, then export
Fix any misheard names while listening to the original audio, rename the speakers, then export as SRT. The file uses standard hh:mm:ss,mmm timecodes, so every caption tool and player reads it.
What the SRT is actually for
SRT is the lowest-common-denominator caption format, which is exactly why it goes everywhere:
- Uploading real captions to YouTube or Vimeo instead of accepting auto-captions
- Importing a caption track into Final Cut Pro, Premiere, DaVinci Resolve, iMovie, or CapCut
- Sitting beside a video file as a sidecar for VLC, Plex, or Infuse
- Interview transcripts where every line has to carry a timecode
- Accessibility requirements that ask for a caption file rather than burned-in text
How long it takes
Captioning a back catalogue is where on-device processing changes the arithmetic — there is no queue and no monthly allowance to spend:
| Recording length | Time to transcribe |
|---|---|
| 5 minutes | about 16 seconds |
| 30 minutes | about 1.5 minutes |
| 1 hour | about 3 minutes |
| 3 hours | about 10 minutes |
Measured by LoroNote on an iPhone 15 Pro at roughly 19x real time. Newer devices transcribe faster; older supported models take longer.
What this is not
It writes SRT, not styled captions
No fonts, colours, positioning, or burn-in. SRT is deliberately plain text plus timecodes; everything visual belongs to your editor.
SRT only — no VTT or ASS export
If a platform insists on WebVTT, converting an SRT is a trivial step in most caption tools and editors, but LoroNote itself exports SRT.
Timings follow the speech, not a style guide
Lines break where the speaking does. If you are held to a strict characters-per-line or minimum-duration standard, plan on a pass through a caption editor afterwards.
Subtitles come out in the language spoken
Whisper Large V3 Turbo transcribes; it is not trained as a translation model. Foreign-language subtitles are a separate translation job on top of the SRT.
Why creators end up here
Two things push people off cloud captioning. The first is the footage itself: rough cuts, client work, and anything under embargo should not be sitting on a transcription vendor’s storage just to produce a caption file. The second is the meter — cloud services price by the minute, and a back catalogue of long videos empties an allowance fast.
On-device generation removes both. The file never leaves the phone or iPad, a paid plan sets no cap on minutes or files, and the job finishes on a plane. What you trade is the collaborative side: this is a single-user tool, so shared caption review still belongs on a platform built for it.
Common questions
Can I generate subtitles without an internet connection?
Yes. Recognition runs on the device with Whisper Large V3 Turbo, so SRT generation works in airplane mode. Nothing about the file is transmitted at any point.
Which files can I generate subtitles from?
MP4 and MOV video, and MP3, WAV, and M4A audio. For video, LoroNote reads the audio track directly — you do not need to export a separate audio file first.
Are speaker names included in the SRT?
Yes. Speakers are identified on the device, and the names you assign are preserved through export along with the timecodes — which is what makes an interview SRT usable as a transcript, not just as captions.
Can it translate the subtitles into another language?
No. Whisper Large V3 Turbo transcribes in the language being spoken and is not trained for translation, so the SRT comes out in the source language. Translating it afterwards is a separate step in another tool.
Can I edit the text before exporting the SRT?
Yes. You can edit the transcript while the original audio plays, with the current line highlighted, so a misheard name gets fixed against the recording rather than from memory. Search and replace handles a term that recurs throughout.
Is there a limit on how many videos I can caption?
No plan cap on minutes or files. Because there is no server doing the work, there is no per-minute cost to meter — the constraints are device storage, battery, and processing time.
Related
- Video to text MP4 and MOV: the audio track is read straight out of the file, no editor export first.
- Timestamped transcripts Sentence-level timecodes, tap-to-replay, and SRT export that keeps the timing.
- WAV to text Field recordings and uncompressed masters — 600 MB an hour that never has to be uploaded.
Caption something you already published
Generate an SRT for a video that currently relies on auto-captions and compare the two line by line.
Download LoroNote