Skip to content

feat(vad): 512 ms silence tail (quiet mode 576 ms) so pauses don't split a sentence (#655) - #664

Merged
rosscado merged 2 commits into
mainfrom
vad-tail-512
Sep 23, 2026
Merged

rosscado merged 2 commits into
mainfrom
vad-tail-512

Conversation

@rosscado

@rosscado rosscado commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Why

The VAD closed a segment after 320 ms of silence, 2.4× shorter than Silero's own default. Spontaneous speech is full of 400–500 ms hesitations. In the #655 pause sweep, the old setting split a two-phrase sentence at a 400 ms pause 56% of the time, and at 500 ms 89% of the time. In production (saypi-api Stream A):

  • a Pi turn averaged ~5 uploads;
  • ~45% of Pi's sub-second clips were mid-sentence cuts.

Each fragment was transcribed without the rest of its sentence, often with no Whisper prompt at all, and scored for end-of-turn in isolation.

This is the second of the two #655 changes the founder approved. The first was the Silero v6 upgrade (#661).

What changes

preset tail before tail after
balanced (every host + dictation) 320 ms 512 ms
highSensitivity (quiet mode) 384 ms 576 ms (keeps its longer-than-balanced invariant)

Nothing else changes: thresholds, minimum speech and pre-pad are the same. The benchmark's latency column is now anchored to the pre-#655 320 ms tail, so its tables stay comparable, and the shipped row is labelled from the live config.

What it buys, and what it costs

AMI spontaneous-speech benchmark, Silero v6, npm run bench:vad:segmentation:

tail thoughts split into 2+ uploads clips per turn short mid-turn cuts extra latency on the turn's final upload
320 ms (before) 22% 2.47 78 –
512 ms (this PR) 16% 2.02 52 +192 ms
640 ms 11% 1.76 40 +320 ms

The cost is stated plainly: +192 ms from the end of speech to the turn's submit. The last clip uploads 192 ms later, and the submit timer counts from that upload, so the delay carries straight through. It doesn't compound. Dictation's live text also arrives in bigger chunks, since a pause now has to reach 512 ms. That's the knee of the curve; 640 ms buys a bit more for another +128 ms. Each clip also carries more trailing silence (about +250 ms of audio per turn), which matters if anything meters on audio duration.

Offsetting it (argued, not measured): about 0.45 fewer uploads per turn (fewer /transcribe round-trips and fewer Jev endpointing calls), and complete thoughts reaching both the ASR and the endpointing model.

I did not also anchor the submit timer to the true end of speech. That would shorten every patience-bound wait and change the endpointing behaviour the server eval is calibrated against, so it should be its own decision.

Measure after release: saypi-api's turn-outcome eval (#505) and Stream A can compare clips per turn, the ≤1 s share and premature-submit rate before and after.

Verification

  • npm test: 3051 pass. The locked preset spec is updated, and the lifecycle spec now references VAD_CONFIGS instead of literals.
  • Layer 3 Chrome: 48/48. The synthetic clips now carry ~1.2 s of trailing silence (was ~0.5 s). Chromium loops the fake-mic file, so at 0.5 s the new tail had zero margin at the loop seam; every clip now closes inside the file with ≥576 ms to spare (checked against the real v6 model).
  • e2e-firefox VAD smoke: pass.

Refs #655

🤖 Generated with Claude Code

https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4

… pauses don't split a sentence (#655)

The VAD closed a segment after 320 ms of silence, 2.4x shorter than Silero's
own default. Spontaneous speech is full of 400-500 ms hesitations, so a Pi
turn averaged ~5 uploads and ~45% of its sub-second clips were mid-sentence
cuts (saypi-api Stream A). Each fragment was transcribed without the rest of
its sentence and scored for end-of-turn on its own.

On the #655 AMI benchmark with Silero v6 (#661), 512 ms takes spoken thoughts
split into several uploads from 22% to 16% and clips per turn from 2.47 to
2.02. Cost: +192 ms on each turn's final upload. 640 ms would reach 11% for
+320 ms; 512 is the knee. Quiet mode keeps its longer-than-balanced tail
(576 ms), since quiet speech dips under the bar more often.

The benchmark's latency column is now anchored to the pre-#655 320 ms tail,
so its tables stay comparable, and the shipped row is labelled from the live
config.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-23T07:37:09.811602Z a3e7095 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a3e7095d1a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/vad/VADConfigs.ts
// of sub-second clips were mid-sentence cuts. On the AMI benchmark (bench/vad) with v6,
// 512 ms takes thoughts split from 22% to 16% and clips per turn from 2.47 to 2.02, for
// +192 ms on each turn's final upload; 640 ms buys 11% for +320 ms. 512 is the knee.
redemptionMs: 512,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Update the benchmark narrative for the new shipped tail

Changing the shipped tail to 512 ms makes the canonical benchmark write-up contradictory: bench/vad/README.md:250-285 still calls 320 ms the “shipped”/“today” configuration and describes latency relative to it, while the updated benchmark now labels 320 ms as pre-#655 and reports the shipped 512 ms policy at +192 ms. This can cause future tuning work to mistake the historical baseline for the live configuration, so update the fragmentation narrative and table labels alongside this value.

AGENTS.md reference: AGENTS.md:L104-L108

Useful? React with 👍 / 👎.

… for the longer VAD tail

Chromium's fake-audio-capture loops the clip, so the silence inside the file
is all the VAD gets before speech restarts. The clips had ~0.5 s, which the
Silero v6 model sees as 14-15 sub-threshold frames: zero margin against the
new 512 ms (16-frame) tail and never enough for quiet mode's 576 ms. The
required e2e passed, but only by the loop seam, a latent flake (found by the
#664 review).

Appended 0.7 s of digital silence to every pool clip and both canonical
copies (no re-synthesis, so the speech is unchanged), and set the
generator's pad to 1.2 s to match. Every clip now closes as one segment
with at least 576 ms to spare inside the file at both tails.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4
@rosscado

rosscado commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor Author

Review: APPROVE WITH NITS (independent reviewer); nits addressed

  • Latent flake (fixed in de20dda): the E2E fake-mic clips had ~0.5 s of trailing silence. Chromium loops the file, so a 512 ms tail had zero margin at the loop seam, and 576 ms never fired. The required e2e passed only by luck of the seam. I appended 0.7 s of digital silence to every clip (the speech is untouched) and set the generator to match. Verified with the real v6 model: every clip now closes as one segment with ≥576 ms to spare inside the file.
  • Latency wording (fixed in the description): the +192 ms lands on the submit, not just the upload, because the timer counts from the VAD end event. It's still +192 ms and doesn't compound.
  • Bench README "today" wording: fixed.
  • Quiet mode 576 ms: forced by the existing invariant that quiet mode's tail exceeds balanced's (+64 ms, as before). Flagged to the founder for a one-line acknowledgement; not blocking.
  • Verified by the reviewer:

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant