Skip to content

feat(vad): upgrade to vad-web 0.0.31 + Silero v6, fed with its context window (#655) - #661

Merged
rosscado merged 3 commits into
mainfrom
vad-web-v6-upgrade
Sep 23, 2026
Merged

rosscado merged 3 commits into
mainfrom
vad-web-v6-upgrade

Conversation

@rosscado

Copy link
Copy Markdown
Contributor

Why

The #655 investigation found that vad-web 0.0.24 never gave the Silero model the 64-sample context window it expects in front of each 512-sample frame. It was fixed upstream in 0.0.31 (ricky0123/vad#263). So we've been running v5 on degraded input, and tuning our thresholds against it.

Upgrading to vad-web 0.0.31 and Silero v6 (fed correctly), with unchanged thresholds and silence tail, gives these results on the benchmarks from #657/#659:

before after
clipped short words (FRR, real corpus) 8% 0%
noise/music opening a segment (FAR) 41% 15%
spoken thoughts split into 2+ uploads (AMI) 30% 22%
clips per conversational turn (AMI) 2.86 2.47
end-of-turn latency – unchanged

Fewer noise-triggered segments also means fewer "Thank you." hallucinations reaching ASR. The founder approved the upgrade on #655. The longer silence tail is a separate follow-up PR, so this one carries no latency cost.

What changed

  • Dependencies:

    • @ricky0123/vad-web 0.0.24 → 0.0.31.
    • onnxruntime-web 1.14 → 1.30. That's the version vad-web resolves, so npm keeps one copy.
  • ORT files: ORT ≥ 1.19 ships one CPU WASM build, ort-wasm-simd-threaded.wasm, plus an .mjs glue it loads with import(). We ship those 2 files instead of 4.

    • copy-onnx-files.js copies exactly those two, and fails the build if either is missing.
    • The extension shrinks from 59 MB to 32 MB.
    • The built bundles now contain no eval( / new Function( (main had 3 per content script).
  • Presets (VADConfigs.ts): they use vad-web's ms options and model: "v6". The values are the old frame counts × 32 ms, so segmentation timing is identical. A new spec invariant keeps every duration on the 32 ms grid, because vad-web floors off-grid values.

    • The unused silero_vad_v5.onnx, silero_vad_legacy.onnx and silero_vad.onnx are gone; silero_vad_v6.onnx is added.
    • The stale 1.19-era public/ort-wasm-simd-threaded.mjs is untracked. It's now copied at build time.
  • Mic ownership (src/vad/micStreamLifecycle.ts): vad-web 0.0.27 moved the mic into the library. pause() stops the tracks and start() re-runs getUserMedia. The conversation pauses the VAD every assistant turn, so that would reopen the mic each turn: first-word clipping, a flickering indicator, and possible Firefox re-prompts. Both clients now:

    • open the mic themselves at initialize, so permission errors still surface there;
    • hold it across pause/resume;
    • release it on destroy.

    start() is async now and is awaited, so an audio-graph failure is reported instead of left as an unhandled rejection.

  • Firefox (in-page VAD): I extended the Firefox smoke to actually start the VAD. It caught two regressions that 0.0.31 would have shipped. Both fail on main + 0.0.31, and both pass with the fixes:

    1. Wrong asset base. OnscreenVADClient pointed at getURL("public/"), a directory the build doesn't have. Only RequestInterceptor's fetch rewrite had been rescuing it, and ORT's import() of its glue can't be rewritten. The fix points it at the extension root, as the offscreen handler already does.
    2. No ScriptProcessor fallback. On main, Firefox's AudioWorklet setup fails (AbortError) and 0.0.24 silently fell back to ScriptProcessor. 0.0.31 no longer falls back, so the in-page client now asks for processorType: "ScriptProcessor" on Firefox. That's what Firefox has always actually run.

    The three init "strategies", which only existed to work around the wrong path, collapse to one. RequestInterceptor's list, which also scopes the Firefox same-realm model shim, now names the v6 model and the new WASM.

  • Admission gate: the fallback bar for the none preset tracks vad-web's default positive threshold, which moved from 0.5 to 0.3. A test pins it to the installed library, so the next upgrade can't drift it silently.

  • Bench harness: ported to 0.0.31. It defaults to what ships, and --no-context replays 0.0.24's feeding. It reproduces the Conversation mode uploads very short clips (31% ≤1s on Pi vs 5% in dictation) — investigate VAD segmentation #655 before/after numbers exactly, which also confirms that vad-web's own context handling matches what was benchmarked.

  • Docs: src/vad/README.md is rewritten (the "why all 4 WASM files" section was now wrong), along with CLAUDE.md's binary assets, the store-permissions evidence, the AMO eval note, the .cursor VAD lines and the e2e comments.

Verification

  • npm test: type-check + Jest + Vitest, 3044 pass. New specs cover:
    • the mic-holding contract (vad_handler-mic-lifecycle.spec.ts);
    • the Firefox asset base and processor choice;
    • the library-default pin;
    • the ms-grid invariant.
  • Layer 3 (Chrome, headless): 48/48. This includes the fake-mic and synthetic-audio VAD → STT → transcript turns, which run the v6 model in the offscreen document under MV3 CSP.
  • e2e-firefox: the smoke now presses the dictation button with Firefox's fake mic, asserts the in-page VAD starts, and lets it run for 3 s on fake-mic frames. Main passes it, and so does this branch.
  • Real-host Layer 4: to follow, as a PR comment.

Refs #655

🤖 Generated with Claude Code

https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4

…t window (#655)

vad-web 0.0.24 never gave the Silero model the 64-sample context window it
expects in front of each 512-sample frame (fixed upstream in 0.0.31,
ricky0123/vad#263), so we have been running v5 on degraded input and tuning
our thresholds against it. On the #655 benchmarks, Silero v6 fed correctly,
at unchanged thresholds and silence tail, takes clipped short words from 8% to
0%, noise/music false-accepts from 41% to 15%, and the share of spoken
thoughts split into several uploads from 30% to 22%. No latency change.

What the upgrade needed:
- onnxruntime-web 1.14 -> 1.30 (the version vad-web resolves; one copy). ORT
  now ships one WASM build plus an .mjs glue loaded with import(), so we ship
  2 ORT files instead of 4 (extension 59 MB -> 32 MB, and no eval() left in
  the bundles). copy-onnx-files.js copies exactly those and fails the build if
  one is missing.
- Presets move to vad-web's ms options with model "v6"; values are the old
  frame counts x 32 ms, so segmentation timing is unchanged. The unused v5,
  legacy and v4 model files go.
- vad-web 0.0.27 took over the mic: pause() stopped the tracks and start()
  re-ran getUserMedia, which would reopen the mic every assistant turn. Both
  VAD clients now open the mic themselves and hold it (micStreamLifecycle),
  keeping 0.0.24's behaviour; start() is awaited so audio-graph failures are
  reported rather than left unhandled.
- Firefox (in-page VAD): the extended Firefox smoke caught two regressions
  before they shipped. The asset base pointed at a public/ dir the build does
  not have (only RequestInterceptor's fetch rewrite had been rescuing it, and
  it cannot rewrite ORT's import()), and 0.0.27+ no longer falls back from
  AudioWorklet to ScriptProcessor, which is what Firefox has always actually
  run. Fixed both, and collapsed the three init "strategies" that only
  existed to work around the wrong path.
- The admission gate's fallback bar tracks vad-web's default positive
  threshold (0.5 -> 0.3), with a test that pins it to the installed library.

Verification: npm test (3044 pass); Layer 3 Chrome E2E 48/48, including the
fake-mic and synthetic VAD -> STT turns; e2e-firefox smoke now starts the
in-page VAD and lets it run on fake-mic frames (fails on main+0.0.31 without
the Firefox fixes, passes with them; main passes too). The #655 benchmark
harness is ported and reproduces its before/after numbers exactly.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-23T07:02:18.988826Z c3f5bea PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c3f5beaf85

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread .gitignore
public/*.js.map
public/*.wasm
# ORT glue copied from node_modules by copy-onnx-files.js at build time (#655)
public/ort-wasm*.mjs

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Copy the ORT glue before launching Firefox dev

On a clean checkout, this ignore rule and the accompanying deletion leave ort-wasm-simd-threaded.mjs absent from public/. The installed npm scripts documentation says pre/post scripts require “a matching name,” so npm run dev:firefox does not run the existing predev; there is no predev:firefox to invoke copy-onnx-files.js. ORT 1.30 then fails its dynamic import when the Firefox VAD initializes. Add a matching lifecycle hook or keep the required glue available.

AGENTS.md reference: AGENTS.md:L84-L88

Useful? React with 👍 / 👎.

…e preset on START, drop stale ORT files

Two independent reviews of the vad-web 0.0.31 upgrade found real problems,
all rooted in 0.0.27 moving work from MicVAD.new() into the first start():

- A failed first start() left vad-web stuck in "initializing": later starts
  resolved as if they worked, the mic stayed held, and destroy() threw
  before releasing the ORT session or AudioContext. A stop/destroy racing
  that first start() was also lost. Both clients now create the AudioContext
  themselves and build the audio graph at initialize (createWarmMicVad: one
  start->pause), as 0.0.24's new() did, so failures surface at init and we
  can always close what we opened. An instance whose start fails is dropped,
  so the next start rebuilds it.
- The `none` preset was reachable in production: after the offscreen
  document's 30 s idle auto-shutdown, the next START re-created it with no
  preset, and under 0.0.27+ library defaults that meant a 1400 ms tail and a
  400 ms minimum (and quiet mode silently dropped). OffscreenVADClient now
  re-sends its init options with START, both clients fall back to
  `balanced`, and the unreachable `none` preset is deleted (#571 rule).
- ORT files in public/ are git-ignored, so a checkout that had built ORT
  1.14 still held its three extra ~9 MB WASM variants and a release build
  would ship them (and differ from AMO's clean rebuild). copy-onnx-files.js
  now removes any ORT file it doesn't copy.
- The in-page client now also runs ORT single-threaded with no proxy (shared
  ortRuntime.ts), so a cross-origin-isolated host can't make ORT spawn a
  blob worker under the page's CSP.
- Docs that contradicted the change (caution map, README, release notes on
  eval) are corrected; the Firefox smoke now requires "Waiting for speech".

Verification: npm test 3050 pass; Layer 3 Chrome 48/48; e2e-firefox smoke
starts the in-page VAD and runs it on fake-mic frames.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4
@rosscado

Copy link
Copy Markdown
Contributor Author

Review round 1: two independent reviewers, both REQUEST CHANGES; addressed in c6c7fb9

Runtime lens:

  1. none preset reachable in production. After the offscreen doc's 30 s idle auto-shutdown, the next preset-less START re-created the VAD on library defaults. On 0.0.27+ that means a 1400 ms tail and a 400 ms minimum.
    → Fixed. OffscreenVADClient re-sends its init options with START. Both clients now fall back to balanced, and none is deleted. Tests: vad_handler-mic-lifecycle.spec.ts (preset on START + fallback) and OffscreenVADClient-startPreset.spec.ts.
  2. A failed first start() left a zombie. vad-web stuck in "initializing", later starts reported success, and the session and AudioContext leaked.
    → Fixed. Both clients now create the AudioContext themselves and build the audio graph at initialize (createWarmMicVad: one start→pause, as 0.0.24's new() did). A failure surfaces at init with everything released, and an instance whose start fails is dropped and rebuilt next time. Both are tested.
  3. Stop or destroy racing the first start().
    → Resolved by the same change. With the graph built at init, a later start() sets listening synchronously and only reconnects the held stream.
  • Nits:
    • Two ORT builds per bundle: deferred. See below.
    • Smoke readiness: fixed. It now requires "Waiting for speech".

Build/packaging lens:

  1. Stale ORT 1.14 WASM in existing checkouts' public/ would ship. That's ~61 MB instead of 32, and a mismatch with AMO's clean rebuild.
    → Fixed. copy-onnx-files.js now removes any ort-wasm* file it doesn't copy. Verified against real stale files.
  2. Docs contradicting the change (caution map, README, src/vad/README numThreads claim, release eval notes).
    → Fixed. The in-page client now does run ORT single-threaded via the shared ortRuntime.ts, which also protects cross-origin-isolated hosts. Tested.
  3. Nits: the leftover "v5" wording is fixed.

Not changed here, for founder follow-up:

  • scripts/release.mjs:538 AMO approvalNotes still mention the onnxruntime eval warning. That file is path-guarded (release machinery). The note is now stale but harmless, because the bundles contain no eval.
  • Two ORT builds per bundle (~370 KB dead code per script). vad-web's CJS index pulls NonRealTimeVAD → full onnxruntime-web. The clean fix is a Vite resolve.alias in wxt.config.ts, which is founder-gated. Even with it, the content scripts are smaller than main's. I'll file an issue.

Verification on c6c7fb9:

  • npm test: 3050 pass.
  • Layer 3: 48/48.
  • e2e-firefox VAD smoke: pass.

…the vestigial default-threshold constant

- initializeVAD's failure path now releases the audio *this* attempt opened
  (a local reference) rather than whatever the module-level vadAudio holds,
  so a destroy + re-init landing mid-init can't have its audio closed by the
  first attempt's catch.
- With the `none` preset gone every preset defines positiveSpeechThreshold,
  so VAD_LIBRARY_DEFAULT_POSITIVE_THRESHOLD (and its library-pin spec) guarded
  a value nothing used. The trackers are seeded from `balanced` and
  SegmentStatsTracker now requires a threshold.
- The preemption spec clears the warm-up pause right after tab 1's start, so
  tab 2's takeover stays covered by the not-paused assertion.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4
@rosscado

Copy link
Copy Markdown
Contributor Author

Review round 2: both reviewers APPROVE WITH NITS; nits addressed in 7fd6968

Runtime lens: APPROVE WITH NITS. Verified against vad-web 0.0.31's source:

  • Warm-up is side-effect free. Only a microtask separates the warm start from the pause. Frames arrive as tasks after the frame processor is inactive, so no callbacks fire in between.
  • No double-connected sources. Start/pause swap the stored source node, and a pause can't land mid-start from a separate message.
  • Everything is released on every path. The AudioContext is ours (ownsAudioContext false) and is closed on init failure, warm failure, start failure and destroy; the mic is released on the same paths.
  • The preset flows end to end: client → offscreen_manager forward → handler, with a balanced fallback.
  • No remaining regression versus 0.0.24.

Nits:

  • The init catch released the module-level vadAudio rather than this attempt's audio. Fixed with a local reference.
  • If the warm-up fails before the graph exists, vad-web's destroy() throws before model.release(), so the ORT session is not released. Accepted: it only happens on an init that has already failed, and fixing it means constructing the Silero model ourselves.

Build/packaging lens: APPROVE WITH NITS. Specs are strengthened, not weakened: the preset keys are pinned, and the invariants now cover every preset. The stale-file removal is safe (top level only, ort-wasm* names not in the copy list, nothing tracked matches).

Nits:

  • The preemption spec's spy-clear masked tab 2's takeover. Fixed: it now clears right after tab 1's start.
  • VAD_LIBRARY_DEFAULT_POSITIVE_THRESHOLD was vestigial once none went. Removed: trackers are seeded from balanced, and SegmentStatsTracker requires a threshold.

Real-host (Layer 4, real Chrome, real STT):

Final local verification on 7fd6968:

  • npm test: 3049 pass.
  • Layer 3: 48/48.
  • e2e-firefox VAD smoke: pass.

Follow-ups filed:

@rosscado
rosscado merged commit 947d5a1 into main Sep 23, 2026
5 checks passed
@rosscado
rosscado deleted the vad-web-v6-upgrade branch September 23, 2026 07:29
rosscado added a commit that referenced this pull request Sep 23, 2026
…lit a sentence (#655) (#664)

* feat(vad): lengthen the silence tail to 512 ms (quiet mode 576 ms) so pauses don't split a sentence (#655)

The VAD closed a segment after 320 ms of silence, 2.4x shorter than Silero's
own default. Spontaneous speech is full of 400-500 ms hesitations, so a Pi
turn averaged ~5 uploads and ~45% of its sub-second clips were mid-sentence
cuts (saypi-api Stream A). Each fragment was transcribed without the rest of
its sentence and scored for end-of-turn on its own.

On the #655 AMI benchmark with Silero v6 (#661), 512 ms takes spoken thoughts
split into several uploads from 22% to 16% and clips per turn from 2.47 to
2.02. Cost: +192 ms on each turn's final upload. 640 ms would reach 11% for
+320 ms; 512 is the knee. Quiet mode keeps its longer-than-balanced tail
(576 ms), since quiet speech dips under the bar more often.

The benchmark's latency column is now anchored to the pre-#655 320 ms tail,
so its tables stay comparable, and the shipped row is labelled from the live
config.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4

* test(e2e): give the synthetic speech clips ~1.2 s of trailing silence for the longer VAD tail

Chromium's fake-audio-capture loops the clip, so the silence inside the file
is all the VAD gets before speech restarts. The clips had ~0.5 s, which the
Silero v6 model sees as 14-15 sub-threshold frames: zero margin against the
new 512 ms (16-frame) tail and never enough for quiet mode's 576 ms. The
required e2e passed, but only by the loop seam, a latent flake (found by the
#664 review).

Appended 0.7 s of digital silence to every pool clip and both canonical
copies (no re-synthesis, so the speech is unchanged), and set the
generator's pad to 1.2 s to match. Every clip now closes as one segment
with at least 576 ms to spare inside the file at both tails.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01W68zjaRTjQGYGCvoBA6aJ4

---------

Co-authored-by: Claude Opus 5.5 (1M context) <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant