Real-time English subtitles for anything playing on your Mac, plus a one-click
export of the whole session to .srt.
Runs in the menu bar, entirely on-device — no accounts, no cloud, and no audio ever leaves the machine.
Grab the latest release — a disk image you drag to Applications:
https://github.com/martinx/LiveSubtitles/releases/latest
- Open LiveSubtitles.dmg (or the zip — that is what the in-app updater downloads) and double-click Live Subtitles. It sees that it is running from a download and offers to install itself into Applications, then relaunches from there. Dragging it across yourself works just as well.
- Launch it. macOS asks for Screen Recording — that is how system audio is read. Grant it under System Settings → Privacy & Security → Screen Recording, then relaunch if you had to change it.
- Press ⌥⌘L (or use the menu bar) to start listening.
The first run also downloads and compiles the speech model — see First run.
Why does it insist on living in Applications? A quarantined app that has not been notarised gets run by macOS from a fresh random directory on every launch ("App Translocation") when it is opened from Downloads or from the disk image. Each launch then looks like a different app, so the Screen Recording grant can never be remembered and the permission prompt keeps coming back no matter how many times the box is ticked. Installing once removes the whole problem — and the released builds are signed and notarised so downloads are not quarantined in the first place.
- macOS 14+ (developed and measured on macOS 27, Apple M3)
- Apple Silicon (the speech models run on the Neural Engine)
- Xcode command line tools (
swift)
git clone https://github.com/martinx/LiveSubtitles.git
cd LiveSubtitles
make install # release build + copy to /Applications| Command | What it does |
|---|---|
make |
release build → build/LiveSubtitles.app |
make run |
build and launch from build/ |
make install |
build and copy to /Applications |
make debug |
debug build — faster to compile, slower to run |
make update |
git pull, rebuild, reinstall |
make check-update |
report whether the remote has new commits |
make icon |
regenerate Resources/AppIcon.icns |
make dist |
zip the app in build/ without rebuilding it |
make dmg |
drag-to-Applications disk image, also without rebuilding |
make package |
build, then both artefacts |
make clean |
remove build/ and .build/ |
Then launch it like any other app — Spotlight (Cmd-Space → Live Subtitles),
Launchpad — or open -a LiveSubtitles.
The Screen Recording grant is tied to the app's signature, so updating in place keeps the permission; if macOS asks again, re-grant it under System Settings → Privacy & Security → Screen Recording.
- Screen Recording permission — required to capture system audio. macOS prompts on first launch.
- The model downloads on first use (433 MB for the 120M models, 594 MB for the
0.6B ones) into
~/Library/Application Support/FluidAudio/Models/. - The first load also compiles the CoreML model, which is why it can take 40–90 s. That is a one-off: later launches load in about 2 s, and the compiled model is cached system-wide.
Click the captions bubble in the menu bar. The first line is the current state, and the rest is the usual set of actions:
| Item | What it does |
|---|---|
| (state) | Listening / Paused / Stopped |
| Start Listening | Loads the model if needed and begins capturing |
| Pause Listening | Stops capturing, keeps the model loaded |
| Stop Listening | Stops capturing and releases the model |
| Clear Captions | Empties the on-screen text and the session transcript |
| Export Transcript… | Saves the session as .srt (or .txt) |
| Copy Transcript | Puts the whole transcript on the clipboard |
| Settings… | Engine, behaviour, appearance and shortcut options |
| Restart Engine | Rebuilds the engine after an engine setting changed |
| Check for Updates… | Asks GitHub for the latest release (see Keeping up to date) |
| How to Use | The first-run walkthrough, any time |
| About Live Subtitles | Version, author, licence and links |
| Quit | Stops capture and exits |
The menu-bar icon reflects the state: a filled bubble while listening, a pause symbol while paused, an outline bubble when stopped.
Three states, because "stop" and "pause" cost different things:
| State | Capture | Model | Coming back |
|---|---|---|---|
| Listening | running | loaded | — |
| Paused | stopped | still loaded | instant |
| Stopped | stopped | released (~600 MB back) | reloads the model, ~2 s |
Pause exists because the common case is stepping away for a minute: paying the model load again would defeat the point. Stop exists because the common case there is "I am not using this right now", where holding 600 MB is the wrong trade.
The transcript is never touched by pausing or stopping — only Clear Captions empties it. You can stop mid-episode, come back, and export everything captured so far.
Whether the app starts listening on launch is a setting (Start listening when the app launches, on by default).
Global, so they work while the video player is frontmost:
| Action | Default |
|---|---|
| Start listening | ⌥⌘L |
| Pause listening | ⌥⌘P |
| Stop listening | ⌥⌘. |
Change them in Settings → Shortcuts: click Record, press a combination containing at least one modifier, or Clear to unbind. Escape cancels, Delete clears. A combination another app already owns will silently not take effect.
These use Carbon's RegisterEventHotKey rather than NSEvent's global monitor, because
monitoring the keyboard that way requires Accessibility permission — this app should not
have to ask for that just so you can pause.
The overlay is click-through and never takes focus, so it can sit over a full-screen video without interrupting playback.
The engine is a setting, not a hard-coded choice. All are English, on-device, and run on the Neural Engine.
| Model | Params | First word | Punctuation | Download |
|---|---|---|---|---|
| Parakeet EOU @160 ms | 120M | 0.63 s | no | 433 MB |
| Parakeet EOU @320 ms | 120M | ~0.7 s | no | 433 MB |
| Parakeet Unified @320 ms (default) | 0.6B | 0.86 s | yes | 594 MB |
| Parakeet Unified @640 ms | 0.6B | ~1.0 s | yes | 594 MB |
| Nemotron @560 ms | 0.6B | 1.50 s | yes | ~600 MB |
Measured on an M3 by playing a known sentence ("Quantum mechanics explains the behaviour of very small particles") and timestamping the app's own output:
| EOU 120M @160 ms | Unified 0.6B @320 ms | Nemotron 0.6B @560 ms | |
|---|---|---|---|
| First word on screen | 0.63 s | 0.86 s | 1.50 s |
| Transcript | "…very small part of" | "…small particles." | "…small particles." |
| Style | run-on lowercase | punctuated, capitalised, numbers normalised | punctuated |
The default is the 0.6B model: ~0.2 s more latency buys correct words plus real punctuation and capitalisation.
Why not the 640 ms tier? It is the same model, not a more accurate one: FluidAudio describes it as "same WER as 320 ms at ~2.5x the RTFx" — it re-encodes far less often. That is an efficiency tier, so it only helps a machine that is struggling for compute. On an M3 the 320 ms tier runs comfortably, so 640 ms just adds 0.32 s of delay for nothing.
Why not Nemotron? Measured 0.64 s slower to first word with no accuracy gain on the test sentence. Worth trying only if a particular show trips up Parakeet.
| Setting | Notes |
|---|---|
| Start listening when the app launches | On by default. Turn off to launch idle and start with a shortcut. |
| Shortcuts | Global Start / Pause / Stop bindings (own tab). |
| Check for updates automatically | Daily GET of the public GitHub releases API. No identifiers are sent. |
| Model | Which streaming speech model to run. Needs Apply & Restart Engine. |
| End of sentence after | How much silence closes a subtitle line. Needs Apply & Restart Engine. |
| New line after silence | Off / 2 / 3 / 5 / 8 s. After that much quiet, the next sentence starts a fresh line instead of being appended to the previous one. |
| Keep captions on screen | Stop the overlay fading out during quiet stretches. |
| Show Dock icon | Off by default — the app is a menu-bar accessory and never takes activation from the video. Turn on for a Dock icon, a Cmd-Tab entry, a full app menu, and a LIVE badge on the Dock tile while capturing. |
| Drag to reposition | On by default: grab the caption bar and put it where you want, and the spot is remembered. While it is on, the overlay takes clicks on the bar itself; turn it off to make the overlay fully click-through. |
| Font size / Width / Background / Distance from bottom | Apply immediately |
| Max lines | 1–5. 1 = current sentence only · 2 = previous line above it · 3+ gives the current sentence two lines and keeps more history. |
Once you have dragged the overlay, the saved position wins over Distance from bottom until you press Reset overlay position.
Worth knowing, because three separate things can clear the screen:
- New line (text dropped). After
New line after silenceof quiet audio, the running line is discarded so the next sentence starts fresh. The break is applied when speech resumes, not during the pause — the last line stays readable while nothing is being said. - Fade out. Unless Keep captions on screen is on, the overlay fades out
(
showsCaption = false, 0.28 s) 6 s after the fade timer starts — i.e. after "new line" delay plus 6 s. With Keep captions on screen it never fades. - Empty box. The rounded background disappears whenever there is no text at all.
Pause detection runs on the audio level, not on the model going quiet: every one of these models keeps emitting hallucinated words through silence, which resets a text-based timer forever. The threshold adapts to the show (audio ~10 dB below the recent peak, with a slow-decaying peak follower).
Export writes SubRip (.srt). Cue bounds come from the audio clock — the moment a
cue's first words appeared and the moment its last new word arrived — so timings track
playback for every model family:
1
00:00:01,760 --> 00:00:04,280
He would rather the meeting is scheduled for tomorrow morning at nine.
2
00:00:07,800 --> 00:00:10,200
I told him to wait outside until the police arrive.
Timeline caveat: the clock starts when capture starts, so cue times are relative to when you launched (or restarted) the engine. Start captions at the beginning of an episode if you want the
.srtto line up with the file.
Measured on an M3 with the default model:
| Metric | Value |
|---|---|
| ScreenCaptureKit buffers | 320 frames = 20 ms, ~54 buffers/s |
| First word of a stream | 0.86 s (0.63 s with EOU @160 ms) |
| Per-word lag once running (EOU @160 ms) | 0.1–0.2 s |
| Model load, warm cache | ~2 s |
| Model load, first ever use | 40–90 s (download + one-off CoreML compile) |
| Audio work per buffer | O(1) — no resampling, one 1.3 KB copy |
Design choices that keep it cheap:
- One ordered consumer. Audio is fed to the model from a single detached task via
an
AsyncStream, so buffers cannot be reordered and FluidAudio is only ever touched from one place. - The capture format is already the model's format. ScreenCaptureKit is asked for 16 kHz mono and delivers 16 kHz mono Float32, so nothing is resampled on the hot path.
- O(1) cue segmentation. A cue ends on a finished sentence when the model has punctuated one, otherwise after 700 ms without new words or 22 words. Line breaks come from audio energy — no per-show VAD to tune.
- Cue text is tidied on the way out: stranded leading marks are dropped, runs of
punctuation collapse to the strongest mark (
..→.,,.→.), a full stop is only appended after an actual word, and cues with no words at all are discarded. - Bounded rendering. Only a trailing window of transcript is drawn, partials are
only pushed when the text actually changed, and cues with no actual words (a stray
?from noise) are dropped. - The engine is never reset between sentences — resetting discards encoder context and can clip the start of the next line.
LIVESUBTITLES_DEBUG=1 make run # trace partials, cues, state, hot keys
LIVESUBTITLES_DUMP_SRT=/tmp/out.srt make run # mirror the transcript to disk
LIVESUBTITLES_OPEN=settings make run # also: welcome, aboutTrace tags: [partial], [utterance], [pause], [newline], [state], [hotkey],
[drag], [status], [cue].
Original artwork, drawn from scratch by scripts/make-icon.py (Pillow): a dark
"screen" carrying a white caption card with two text lines and a speech tail, plus a
coral dot for "live". No third-party icon assets are used.
python3 scripts/make-icon.py # writes icon_1024.png + a size-legibility strip
make build # bundles Resources/AppIcon.icnsPackages/LiveSubtitlesKit/ storage and (later) analysis — no AppKit, tested on its own
Sources/LiveSubtitles/
├── App.swift entry point (accessory unless the Dock icon is on)
├── AppDelegate.swift status item, status menu, app main menu, Dock state
├── GlobalHotKeyCenter.swift system-wide hot keys (Carbon)
├── KeyShortcut.swift a shortcut + its label and Carbon modifier mask
├── ShortcutRecorder.swift click-to-record shortcut field
├── CaptionController.swift wires everything; owns export + settings window
├── SystemAudioCapture.swift ScreenCaptureKit -> AsyncStream<AVAudioPCMBuffer>
├── StreamingTranscriber.swift multi-model streaming ASR + cue segmentation + pauses
├── TranscriptStore.swift timestamped cues -> .srt / .txt
├── CaptionModel.swift committed vs in-progress text, fade-out
├── CaptionView.swift rolling caption, finished and live text on separate lines
├── CaptionPanel.swift non-activating, click-through, draggable overlay window
├── Settings.swift UserDefaults-backed settings (+ migration)
└── SettingsWindow.swift grouped-form settings UI
Data flow:
ScreenCaptureKit (16 kHz mono Float32, 20 ms buffers)
│
▼
StreamingTranscriber ── selected model (EOU / Unified / Nemotron) on the ANE
│ partial text │ cue bounds from the audio clock
│ audio-level pauses │
▼ ▼
CaptionModel TranscriptStore
│ │
▼ ▼
CaptionPanel .srt / .txt export
This began as a rewrite of an app whose macOS target used SwiftFasterWhisper. That
pipeline buffered a fixed 4.2 s window (WINDOW_SIZE_SAMPLES = 67200) before every
decode, which put a hard ~4–6 s floor under subtitle latency that no model or parameter
change could remove. It also needed a 1.5 GB model and a CTranslate2 framework whose
headers were pruned, so it would not even compile without patching the dependency.
- English only.
- No punctuation from the EOU models. The default 0.6B model punctuates and capitalises; switch to EOU and you get run-on lowercase with a full stop appended.
- A committed cue is frozen. Streaming models keep revising the sentence they are on, but once a cue has been cut and added to the transcript, later revisions are not written back into it.
- No automatic recovery for the capture stream. If ScreenCaptureKit stops (display change, permission change, monitor unplugged) audio goes dead until you pick Restart Engine.
- Word errors and occasional hallucination on music-heavy or heavily accented audio.
- Offline transcription of local files. Point the app at an episode, transcribe it
ahead of time with a batch model (Parakeet TDT v2 / Ultra) and export a perfectly
timed
.srt— zero latency and higher accuracy than any streaming pass, because the whole file is available up front. Streaming stays for anything that cannot be pre-processed. - Auto-reconnect the capture stream so a display change does not need a manual restart.
The app knows its own version and can update itself:
- Automatically — once a day it asks the public GitHub releases API for the latest
tag, if Check for updates automatically is on (Settings → General). Nothing about
you is sent; it is a plain
GET. - On demand — Check for Updates… in the menu.
- Installing — when there is a newer release, the menu offers Update to x.y.z….
It downloads the zip, unpacks it, checks that the bundle reports the expected version
and passes
codesign --verify, then hands the swap to a short script that waits for the app to exit, keeps a backup of the old bundle, and rolls back if the swap fails. That is deliberately not silent: the app carries a Screen Recording grant, and silently replacing it is a good way to end up with something that will not launch. - Self-install only works from
/Applications. Anywhere else the app says so rather than guessing.
If you built from source, make update pulls the latest commit, rebuilds and reinstalls.
Everything is on-device. There are exactly two network requests in the whole app:
- The speech model download — one time, from Hugging Face, on first use.
- The update check — a
GETof the public GitHub releases API, at most once a day, and it can be switched off.
No audio, no transcript, no identifiers, no analytics, no account.
Tag and push; the release workflow builds the bundle,
stamps the version from the tag, and publishes a GitHub Release with
LiveSubtitles.zip attached:
git tag v0.2.0
git push origin v0.2.0Without any secrets that produces a working but ad-hoc signed build: it runs, but on someone else's Mac Gatekeeper blocks the first launch until they right-click → Open. Signing with a Developer ID Application certificate and notarising removes that warning completely.
You need a paid Apple Developer Program membership (a Developer ID certificate cannot be issued to a free team) and an App Store Connect API key.
-
Create an API key. appstoreconnect.apple.com → Users and Access → Integrations → App Store Connect API →
+, role Admin or App Manager. Download the.p8(Apple only lets you download it once) and note the Key ID; the Issuer ID is shown above the key list. -
Run one command:
scripts/setup-signing.sh ~/Downloads/AuthKey_XXXXXXXXXX.p8It generates the private key and CSR locally, asks Apple to issue a Developer ID Application certificate for it, builds the
.p12, imports it into your login keychain so local builds sign too, and uploads every secret the workflow needs. -
Prove it works without publishing:
gh workflow run release.yml -f tag=v0.1.2 -f dry_run=true
The whole path runs — import the certificate, sign with hardened runtime and a secure timestamp, notarise, staple — and the artefacts are uploaded as workflow artefacts instead of a release. If Import the Developer ID certificate, Notarise and Check Gatekeeper's verdict are green, the next tag produces a download that just opens.
Export it with its private key — Keychain Access → login → My Certificates →
right-click Developer ID Application: … → Export… → .p12, set a password — then:
scripts/make-signing-secrets.sh DeveloperID.p12 AuthKey_XXXXXXXXXX.p8An APPLE_ID + APPLE_APP_PASSWORD + APPLE_TEAM_ID triple also works if you would
rather use an app-specific password than an API key. Every signing step is conditional,
so the workflow stays green with or without any of this.
Whose name appears? A Developer ID is issued to a team. Users see that team's name in the Gatekeeper prompt — for example "…was signed by Shanghai Dst Technology Co.,ltd.". For a personal project under your own name you would need a separate, personally enrolled paid account.
The landing page in docs/ is published by GitHub Pages at
https://subtitles.bitey.ai.
MIT. Bundled third-party components and the model licences are set out in THIRD_PARTY_NOTICES.md — notably FluidAudio, which is Apache 2.0 and is linked statically into the binary.
Bug reports and focused fixes are welcome — see CONTRIBUTING.md. The one hard rule: everything stays on-device.
