feat(vad): 512 ms silence tail (quiet mode 576 ms) so pauses don't split a sentence (#655) - #664
Conversation
… pauses don't split a sentence (#655) The VAD closed a segment after 320 ms of silence, 2.4x shorter than Silero's own default. Spontaneous speech is full of 400-500 ms hesitations, so a Pi turn averaged ~5 uploads and ~45% of its sub-second clips were mid-sentence cuts (saypi-api Stream A). Each fragment was transcribed without the rest of its sentence and scored for end-of-turn on its own. On the #655 AMI benchmark with Silero v6 (#661), 512 ms takes spoken thoughts split into several uploads from 22% to 16% and clips per turn from 2.47 to 2.02. Cost: +192 ms on each turn's final upload. 640 ms would reach 11% for +320 ms; 512 is the knee. Quiet mode keeps its longer-than-balanced tail (576 ms), since quiet speech dips under the bar more often. The benchmark's latency column is now anchored to the pre-#655 320 ms tail, so its tables stay comparable, and the shipped row is labelled from the live config. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a3e7095d1a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| // of sub-second clips were mid-sentence cuts. On the AMI benchmark (bench/vad) with v6, | ||
| // 512 ms takes thoughts split from 22% to 16% and clips per turn from 2.47 to 2.02, for | ||
| // +192 ms on each turn's final upload; 640 ms buys 11% for +320 ms. 512 is the knee. | ||
| redemptionMs: 512, |
There was a problem hiding this comment.
Update the benchmark narrative for the new shipped tail
Changing the shipped tail to 512 ms makes the canonical benchmark write-up contradictory: bench/vad/README.md:250-285 still calls 320 ms the “shipped”/“today” configuration and describes latency relative to it, while the updated benchmark now labels 320 ms as pre-#655 and reports the shipped 512 ms policy at +192 ms. This can cause future tuning work to mistake the historical baseline for the live configuration, so update the fragmentation narrative and table labels alongside this value.
AGENTS.md reference: AGENTS.md:L104-L108
Useful? React with 👍 / 👎.
… for the longer VAD tail Chromium's fake-audio-capture loops the clip, so the silence inside the file is all the VAD gets before speech restarts. The clips had ~0.5 s, which the Silero v6 model sees as 14-15 sub-threshold frames: zero margin against the new 512 ms (16-frame) tail and never enough for quiet mode's 576 ms. The required e2e passed, but only by the loop seam, a latent flake (found by the #664 review). Appended 0.7 s of digital silence to every pool clip and both canonical copies (no re-synthesis, so the speech is unchanged), and set the generator's pad to 1.2 s to match. Every clip now closes as one segment with at least 576 ms to spare inside the file at both tails. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4
Review: APPROVE WITH NITS (independent reviewer); nits addressed
|
Why
The VAD closed a segment after 320 ms of silence, 2.4× shorter than Silero's own default. Spontaneous speech is full of 400–500 ms hesitations. In the #655 pause sweep, the old setting split a two-phrase sentence at a 400 ms pause 56% of the time, and at 500 ms 89% of the time. In production (saypi-api Stream A):
Each fragment was transcribed without the rest of its sentence, often with no Whisper prompt at all, and scored for end-of-turn in isolation.
This is the second of the two #655 changes the founder approved. The first was the Silero v6 upgrade (#661).
What changes
balanced(every host + dictation)highSensitivity(quiet mode)Nothing else changes: thresholds, minimum speech and pre-pad are the same. The benchmark's latency column is now anchored to the pre-#655 320 ms tail, so its tables stay comparable, and the shipped row is labelled from the live config.
What it buys, and what it costs
AMI spontaneous-speech benchmark, Silero v6,
npm run bench:vad:segmentation:The cost is stated plainly: +192 ms from the end of speech to the turn's submit. The last clip uploads 192 ms later, and the submit timer counts from that upload, so the delay carries straight through. It doesn't compound. Dictation's live text also arrives in bigger chunks, since a pause now has to reach 512 ms. That's the knee of the curve; 640 ms buys a bit more for another +128 ms. Each clip also carries more trailing silence (about +250 ms of audio per turn), which matters if anything meters on audio duration.
Offsetting it (argued, not measured): about 0.45 fewer uploads per turn (fewer /transcribe round-trips and fewer Jev endpointing calls), and complete thoughts reaching both the ASR and the endpointing model.
I did not also anchor the submit timer to the true end of speech. That would shorten every patience-bound wait and change the endpointing behaviour the server eval is calibrated against, so it should be its own decision.
Measure after release: saypi-api's turn-outcome eval (#505) and Stream A can compare clips per turn, the ≤1 s share and premature-submit rate before and after.
Verification
npm test: 3051 pass. The locked preset spec is updated, and the lifecycle spec now referencesVAD_CONFIGSinstead of literals.Refs #655
🤖 Generated with Claude Code
https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4