A local voice-over studio for Apple Silicon, built on VoxCPM2. Write a script, pick or clone a voice, render — everything runs on your Mac, nothing leaves it.
Doublure is French for a stand-in or voice double: someone who speaks in your place, as you.
- Voices — design a voice from a text description, clone one from a 10–15 s recording made right in the browser (level meter, playback, optional background-noise reduction), or import a file in any format. Designed voices are frozen into a reference clip so they stay identical across sessions; recordings are auto-transcribed, and the script you read becomes the exact transcript. Each voice has a one-click test, and an optional ultimate cloning mode uses the transcript for higher-fidelity continuation cloning (~40 % slower).
- Studio — a script is a list of paragraphs. Edit one, re-render one, re-roll a take you don't like (per-paragraph seed). Every chunk is cached by content hash, so re-rendering a 40-paragraph script after fixing a typo regenerates one paragraph, not forty.
- Quality gate — each generated chunk is transcribed by Whisper and compared with the input text; if too many words are wrong it is re-rolled with another seed (a take) and the best take is kept. The Studio shows, word by word, what was written versus what Whisper heard.
- Pronunciation fixes — click a mis-heard word in that diff and give the voice a respelling, e.g. Cibelli → Tchibelli. The fix belongs to that paragraph only, and only what is spoken changes: the written text, the SRT and the quality check keep the original word. The paragraph shows up as not rendered until you generate it again.
- Delivery control — a per-project or per-paragraph instruction such as enthusiastic, slightly faster changes how a cloned voice speaks without changing who is speaking; click into the field for suggestions in the script's language. Non-verbal tags go straight into the text: click into a paragraph and the tags VoxCPM2 documents (
[laughing],[sigh],[Uhm],[Shh]) appear under it, one click inserts one at the cursor; tags are highlighted in the text and an unknown one such as[rire]is flagged with the tag to use instead. Punctuation shapes pauses (periods, commas, ellipsis) more than intensity. - Guidance built in — a guide panel (header
?, or the?key) holds every rule in one scroll: script syntax, tags, delivery, takes and seeds, quality gate, fit to picture, recording a good clip. Short contextual tips appear once, when they apply (first render, first quality warning, first paragraph over budget, several voices in the library…), and can be brought back from the sidebar. The empty Projects page offers an example project: a short script that uses every trick, to read and then generate. - Fit to picture — attach a video (or import its SRT, or type times by hand) and every paragraph gets a time window. With a video and no script, Whisper listens to it and cuts one cued paragraph per sentence. The Studio shows each take against its budget; Fit re-takes a paragraph that runs long — other seeds first, then progressively faster delivery — and keeps the shortest take that passes the quality gate. Fit all does it for the whole script.
- Dialogue scripts — start a paragraph with a voice name and a colon (
NARRATEUR: …) and it gets that voice; any paragraph can also pick its voice in the Studio. One project, one export, several speakers. - Export — WAV normalized to −16 LUFS / −1 dBTP, MP3, and sentence-level SRT. Cued paragraphs are placed at their time on the timeline (never overlapping); with a video attached, an MP4 with the new voice track is written as well.
- Draft / Final — 5 or 10 diffusion steps; draft is ~2× faster for iterating on wording.
- Queue — job history with links back to the project or voice, per-job output files (play, download, reveal in Finder, delete). Voice reference clips are protected; exports can also be revealed straight from the Studio.
Generation is slower than real time (see Performance), so the app is built as a render queue: you queue, you come back.
Three language settings, deliberately independent: the UI (header switch, English default; web/src/lib/i18n.tsx holds EN and FR side by side — add a language by adding a column), the clone script (FR/EN toggle in the clone card — the language you read aloud), and each project's language (drives the Whisper quality check and number spelling; 12 languages offered).
- macOS on Apple Silicon (developed on an M1 Pro, 16 GB)
- Python 3.11 or 3.12 and uv
- Node.js 20+
ffmpeg(loudness normalization, MP3, decoding browser recordings)
brew install uv ffmpeg nodegit clone <this repo> doublure && cd doublure
uv sync # Python env + dependencies
(cd web && npm install) # Next.js UIModels download automatically on first use into the Hugging Face cache (~/.cache/huggingface):
mlx-community/VoxCPM2-8bit (~3.2 GB) when the worker starts, mlx-community/whisper-large-v3-turbo (~1.6 GB) on the first quality check.
./dev.shstarts the three processes (Ctrl-C stops all of them; leftovers from a previous run are stopped first), then open http://localhost:3000.
| process | what | port |
|---|---|---|
doublure worker |
holds the model, drains the job queue | — |
doublure serve |
FastAPI backend | 8787 |
npm run dev (in web/) |
Next.js UI, proxies /api to the backend |
3000 |
The header badge tells you whether the worker is up. If it says worker offline, start .venv/bin/doublure worker in a terminal.
- Voices tab — clone your voice: enter a name, read the on-screen script aloud, listen back, create. A test sentence is generated automatically. Or design one: describe a voice, generate three candidates, pick the best and freeze it.
- Projects tab — paste a script (blank line = new paragraph), choose the voice, create the project. Or click Create the example project to start from a script that shows every trick.
- Studio — generate one paragraph, listen, take another one with New take, then generate all. Optionally set a delivery instruction (project-wide or per paragraph) such as enthusiastic, slightly faster. Then export WAV + MP3 + SRT. Press
?at any time for the guide.
Everything is also scriptable. The worker must be running.
.venv/bin/doublure design "Deep, calm male voice, documentary narrator" --n 3
.venv/bin/doublure voice freeze <job-id> --seed 2 --name Narrator
.venv/bin/doublure voice add Vincent take.m4a --text "exact transcript" --denoise 0.8
.venv/bin/doublure say "Hello there." -v Narrator -o hello.wav [--lang en] [--draft]
.venv/bin/doublure render script.txt -v Narrator -o out.wav [--draft] [--style "enthusiastic, slightly faster"]
.venv/bin/doublure jobsAll optional, via environment variables:
| variable | default | |
|---|---|---|
DOUBLURE_HOME |
./storage |
where the DB, cache, voices and exports live |
DOUBLURE_ENGINE |
mlx |
mlx or llamacpp (see below) |
DOUBLURE_MLX_MODEL |
mlx-community/VoxCPM2-8bit |
-4bit is smaller, -bf16 is larger |
DOUBLURE_MAX_CHUNK_CHARS |
350 |
~15 s of speech; memory grows with chunk length |
DOUBLURE_LANG |
en |
default language for new projects and sample texts; each project has its own language setting that drives the Whisper quality check and number spelling |
DOUBLURE_WHISPER_MODEL |
mlx-community/whisper-large-v3-turbo |
|
DOUBLURE_QC |
1 |
0 disables the quality gate |
DOUBLURE_QC_MAX_WER |
0.12 |
word error rate above which a chunk is re-rolled |
DOUBLURE_QC_RETRIES |
2 |
|
DOUBLURE_FIT_MAX_TRIES |
6 |
takes tried by Fit before keeping the shortest |
DOUBLURE_MAX_UPLOAD_MB |
2048 |
largest video you can attach (checked in the browser, the proxy and the API) |
A lower-memory fallback (flat ~5 GB instead of up to ~9.5 GB, but ~1.7× slower and without transcript-guided cloning):
brew install cmake
git clone --depth 1 https://github.com/tc-mb/llama.cpp-omni.git engines/llama.cpp-omni
cmake -S engines/llama.cpp-omni -B engines/llama.cpp-omni/build -DCMAKE_BUILD_TYPE=Release
cmake --build engines/llama.cpp-omni/build --target voxcpm2-cli -j
# weights from https://huggingface.co/DennisHuang648/VoxCPM2-GGUF into models/
DOUBLURE_ENGINE=llamacpp .venv/bin/doublure workerMeasured on an M1 Pro (16 GB), MLX 8-bit, RTF = generation time / audio duration:
| 10 steps (final) | 5 steps (draft) | |
|---|---|---|
| plain TTS | ~1.6 | ~1.0 |
| cloning (reference clip) | ~1.7–2.1 | ~1.1 |
| ultimate cloning (+ transcript) | ~3.0 | — |
So a 5-minute voice-over is roughly 9 minutes at final quality with a cloned voice. Whisper QC adds ~15 %. Generation is bit-exact for a given seed, which is what makes the cache and frozen voices sound.
doublure/ Python package
engine/ MLX (default) and llama.cpp adapters behind one interface
worker.py resident model process, job queue, per-chunk cache, QC loop
api.py FastAPI backend (never loads the model)
audio.py trimming, joining, loudness, denoise, MP3, SRT
qc.py Whisper transcription + WER
chunker.py paragraph → sentence-packed chunks
cli.py
web/ Next.js 16 UI (App Router, Tailwind)
tests/
bench/ Phase-0 benchmark and probe scripts
storage/ doublure.db, cache/, voices/, renders/, projects/ (attached videos) (gitignored)
Run the tests with .venv/bin/pytest.
- VoxCPM2 by OpenBMB (Apache-2.0) — the speech model
- mlx-audio and mlx-whisper — Apple Silicon inference
- llama.cpp-omni — optional GGUF engine
Voice cloning reproduces a real person's voice. Use it on your own voice or with the explicit consent of the person you're cloning.
MIT — see LICENSE.
