feat(vad): upgrade to vad-web 0.0.31 + Silero v6, fed with its context window (#655) - #661
Conversation
…t window (#655) vad-web 0.0.24 never gave the Silero model the 64-sample context window it expects in front of each 512-sample frame (fixed upstream in 0.0.31, ricky0123/vad#263), so we have been running v5 on degraded input and tuning our thresholds against it. On the #655 benchmarks, Silero v6 fed correctly, at unchanged thresholds and silence tail, takes clipped short words from 8% to 0%, noise/music false-accepts from 41% to 15%, and the share of spoken thoughts split into several uploads from 30% to 22%. No latency change. What the upgrade needed: - onnxruntime-web 1.14 -> 1.30 (the version vad-web resolves; one copy). ORT now ships one WASM build plus an .mjs glue loaded with import(), so we ship 2 ORT files instead of 4 (extension 59 MB -> 32 MB, and no eval() left in the bundles). copy-onnx-files.js copies exactly those and fails the build if one is missing. - Presets move to vad-web's ms options with model "v6"; values are the old frame counts x 32 ms, so segmentation timing is unchanged. The unused v5, legacy and v4 model files go. - vad-web 0.0.27 took over the mic: pause() stopped the tracks and start() re-ran getUserMedia, which would reopen the mic every assistant turn. Both VAD clients now open the mic themselves and hold it (micStreamLifecycle), keeping 0.0.24's behaviour; start() is awaited so audio-graph failures are reported rather than left unhandled. - Firefox (in-page VAD): the extended Firefox smoke caught two regressions before they shipped. The asset base pointed at a public/ dir the build does not have (only RequestInterceptor's fetch rewrite had been rescuing it, and it cannot rewrite ORT's import()), and 0.0.27+ no longer falls back from AudioWorklet to ScriptProcessor, which is what Firefox has always actually run. Fixed both, and collapsed the three init "strategies" that only existed to work around the wrong path. - The admission gate's fallback bar tracks vad-web's default positive threshold (0.5 -> 0.3), with a test that pins it to the installed library. Verification: npm test (3044 pass); Layer 3 Chrome E2E 48/48, including the fake-mic and synthetic VAD -> STT turns; e2e-firefox smoke now starts the in-page VAD and lets it run on fake-mic frames (fails on main+0.0.31 without the Firefox fixes, passes with them; main passes too). The #655 benchmark harness is ported and reproduces its before/after numbers exactly. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c3f5beaf85
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| public/*.js.map | ||
| public/*.wasm | ||
| # ORT glue copied from node_modules by copy-onnx-files.js at build time (#655) | ||
| public/ort-wasm*.mjs |
There was a problem hiding this comment.
Copy the ORT glue before launching Firefox dev
On a clean checkout, this ignore rule and the accompanying deletion leave ort-wasm-simd-threaded.mjs absent from public/. The installed npm scripts documentation says pre/post scripts require “a matching name,” so npm run dev:firefox does not run the existing predev; there is no predev:firefox to invoke copy-onnx-files.js. ORT 1.30 then fails its dynamic import when the Firefox VAD initializes. Add a matching lifecycle hook or keep the required glue available.
AGENTS.md reference: AGENTS.md:L84-L88
Useful? React with 👍 / 👎.
…e preset on START, drop stale ORT files Two independent reviews of the vad-web 0.0.31 upgrade found real problems, all rooted in 0.0.27 moving work from MicVAD.new() into the first start(): - A failed first start() left vad-web stuck in "initializing": later starts resolved as if they worked, the mic stayed held, and destroy() threw before releasing the ORT session or AudioContext. A stop/destroy racing that first start() was also lost. Both clients now create the AudioContext themselves and build the audio graph at initialize (createWarmMicVad: one start->pause), as 0.0.24's new() did, so failures surface at init and we can always close what we opened. An instance whose start fails is dropped, so the next start rebuilds it. - The `none` preset was reachable in production: after the offscreen document's 30 s idle auto-shutdown, the next START re-created it with no preset, and under 0.0.27+ library defaults that meant a 1400 ms tail and a 400 ms minimum (and quiet mode silently dropped). OffscreenVADClient now re-sends its init options with START, both clients fall back to `balanced`, and the unreachable `none` preset is deleted (#571 rule). - ORT files in public/ are git-ignored, so a checkout that had built ORT 1.14 still held its three extra ~9 MB WASM variants and a release build would ship them (and differ from AMO's clean rebuild). copy-onnx-files.js now removes any ORT file it doesn't copy. - The in-page client now also runs ORT single-threaded with no proxy (shared ortRuntime.ts), so a cross-origin-isolated host can't make ORT spawn a blob worker under the page's CSP. - Docs that contradicted the change (caution map, README, release notes on eval) are corrected; the Firefox smoke now requires "Waiting for speech". Verification: npm test 3050 pass; Layer 3 Chrome 48/48; e2e-firefox smoke starts the in-page VAD and runs it on fake-mic frames. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4
Review round 1: two independent reviewers, both REQUEST CHANGES; addressed in c6c7fb9Runtime lens:
Build/packaging lens:
Not changed here, for founder follow-up:
Verification on c6c7fb9:
|
…the vestigial default-threshold constant - initializeVAD's failure path now releases the audio *this* attempt opened (a local reference) rather than whatever the module-level vadAudio holds, so a destroy + re-init landing mid-init can't have its audio closed by the first attempt's catch. - With the `none` preset gone every preset defines positiveSpeechThreshold, so VAD_LIBRARY_DEFAULT_POSITIVE_THRESHOLD (and its library-pin spec) guarded a value nothing used. The trackers are seeded from `balanced` and SegmentStatsTracker now requires a threshold. - The preemption spec clears the warm-up pause right after tab 1's start, so tab 2's takeover stays covered by the not-paused assertion. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4
Review round 2: both reviewers APPROVE WITH NITS; nits addressed in 7fd6968Runtime lens: APPROVE WITH NITS. Verified against vad-web 0.0.31's source:
Nits:
Build/packaging lens: APPROVE WITH NITS. Specs are strengthened, not weakened: the preset keys are pinned, and the invariants now cover every preset. The stale-file removal is safe (top level only, Nits:
Real-host (Layer 4, real Chrome, real STT):
Final local verification on 7fd6968:
Follow-ups filed:
|
…lit a sentence (#655) (#664) * feat(vad): lengthen the silence tail to 512 ms (quiet mode 576 ms) so pauses don't split a sentence (#655) The VAD closed a segment after 320 ms of silence, 2.4x shorter than Silero's own default. Spontaneous speech is full of 400-500 ms hesitations, so a Pi turn averaged ~5 uploads and ~45% of its sub-second clips were mid-sentence cuts (saypi-api Stream A). Each fragment was transcribed without the rest of its sentence and scored for end-of-turn on its own. On the #655 AMI benchmark with Silero v6 (#661), 512 ms takes spoken thoughts split into several uploads from 22% to 16% and clips per turn from 2.47 to 2.02. Cost: +192 ms on each turn's final upload. 640 ms would reach 11% for +320 ms; 512 is the knee. Quiet mode keeps its longer-than-balanced tail (576 ms), since quiet speech dips under the bar more often. The benchmark's latency column is now anchored to the pre-#655 320 ms tail, so its tables stay comparable, and the shipped row is labelled from the live config. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4 * test(e2e): give the synthetic speech clips ~1.2 s of trailing silence for the longer VAD tail Chromium's fake-audio-capture loops the clip, so the silence inside the file is all the VAD gets before speech restarts. The clips had ~0.5 s, which the Silero v6 model sees as 14-15 sub-threshold frames: zero margin against the new 512 ms (16-frame) tail and never enough for quiet mode's 576 ms. The required e2e passed, but only by the loop seam, a latent flake (found by the #664 review). Appended 0.7 s of digital silence to every pool clip and both canonical copies (no re-synthesis, so the speech is unchanged), and set the generator's pad to 1.2 s to match. Every clip now closes as one segment with at least 576 ms to spare inside the file at both tails. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4 --------- Co-authored-by: Claude Opus 5.5 (1M context) <[email protected]>
Why
The #655 investigation found that vad-web 0.0.24 never gave the Silero model the 64-sample context window it expects in front of each 512-sample frame. It was fixed upstream in 0.0.31 (ricky0123/vad#263). So we've been running v5 on degraded input, and tuning our thresholds against it.
Upgrading to vad-web 0.0.31 and Silero v6 (fed correctly), with unchanged thresholds and silence tail, gives these results on the benchmarks from #657/#659:
Fewer noise-triggered segments also means fewer "Thank you." hallucinations reaching ASR. The founder approved the upgrade on #655. The longer silence tail is a separate follow-up PR, so this one carries no latency cost.
What changed
Dependencies:
@ricky0123/vad-web0.0.24 → 0.0.31.onnxruntime-web1.14 → 1.30. That's the version vad-web resolves, so npm keeps one copy.ORT files: ORT ≥ 1.19 ships one CPU WASM build,
ort-wasm-simd-threaded.wasm, plus an.mjsglue it loads withimport(). We ship those 2 files instead of 4.copy-onnx-files.jscopies exactly those two, and fails the build if either is missing.eval(/new Function((main had 3 per content script).Presets (
VADConfigs.ts): they use vad-web's ms options andmodel: "v6". The values are the old frame counts × 32 ms, so segmentation timing is identical. A new spec invariant keeps every duration on the 32 ms grid, because vad-web floors off-grid values.silero_vad_v5.onnx,silero_vad_legacy.onnxandsilero_vad.onnxare gone;silero_vad_v6.onnxis added.public/ort-wasm-simd-threaded.mjsis untracked. It's now copied at build time.Mic ownership (
src/vad/micStreamLifecycle.ts): vad-web 0.0.27 moved the mic into the library.pause()stops the tracks andstart()re-runsgetUserMedia. The conversation pauses the VAD every assistant turn, so that would reopen the mic each turn: first-word clipping, a flickering indicator, and possible Firefox re-prompts. Both clients now:start()is async now and is awaited, so an audio-graph failure is reported instead of left as an unhandled rejection.Firefox (in-page VAD): I extended the Firefox smoke to actually start the VAD. It caught two regressions that 0.0.31 would have shipped. Both fail on main + 0.0.31, and both pass with the fixes:
OnscreenVADClientpointed atgetURL("public/"), a directory the build doesn't have. OnlyRequestInterceptor's fetch rewrite had been rescuing it, and ORT'simport()of its glue can't be rewritten. The fix points it at the extension root, as the offscreen handler already does.AbortError) and 0.0.24 silently fell back to ScriptProcessor. 0.0.31 no longer falls back, so the in-page client now asks forprocessorType: "ScriptProcessor"on Firefox. That's what Firefox has always actually run.The three init "strategies", which only existed to work around the wrong path, collapse to one.
RequestInterceptor's list, which also scopes the Firefox same-realm model shim, now names the v6 model and the new WASM.Admission gate: the fallback bar for the
nonepreset tracks vad-web's default positive threshold, which moved from 0.5 to 0.3. A test pins it to the installed library, so the next upgrade can't drift it silently.Bench harness: ported to 0.0.31. It defaults to what ships, and
--no-contextreplays 0.0.24's feeding. It reproduces the Conversation mode uploads very short clips (31% ≤1s on Pi vs 5% in dictation) — investigate VAD segmentation #655 before/after numbers exactly, which also confirms that vad-web's own context handling matches what was benchmarked.Docs:
src/vad/README.mdis rewritten (the "why all 4 WASM files" section was now wrong), along with CLAUDE.md's binary assets, the store-permissions evidence, the AMO eval note, the.cursorVAD lines and the e2e comments.Verification
npm test: type-check + Jest + Vitest, 3044 pass. New specs cover:vad_handler-mic-lifecycle.spec.ts);Refs #655
🤖 Generated with Claude Code
https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4