Skip to content

Two-way voice: barge-in gated at 150 dB floor + spoken-reply handback - #3

Merged
nastylex merged 2 commits into
mainfrom
rebrand/sirgent-ai
Sep 28, 2026
Merged

nastylex merged 2 commits into
mainfrom
rebrand/sirgent-ai

Conversation

@freebuff-web

@freebuff-web freebuff-web Bot commented Sep 28, 2026

Copy link
Copy Markdown

Summary

Extends the speech plugin (PR #2, merged) with two-way voice conversation:

  • Background playback + mic monitoring — the Stop hook plays the response in the background while watching the microphone (via sox, arecord, or ffmpeg).
  • Barge-in gated at a 150 dB-equivalent floor — SPEECH_BARGE_IN_DB (default 150) is deliberately extreme: ordinary background noise (music, fans, typing, normal conversation ≈ 100–140 dB-eq) can never interrupt; only a sustained near-full-scale signal (~180 dB-eq, a close-range shout held ~2 s) does. Tunable via env var.
  • Spoken-reply handback — a confirmed barge-in stops playback, records up to SPEECH_REPLY_SECS (12 s) of the user's reply, transcribes it with ElevenLabs scribe_v1 STT, and feeds the words back to SirGent via the Stop-hook decision: block contract, so the conversation continues by voice.
  • /voice command — toggles the mode via ~/.sirgent/speech-conversation flag file (wins over SPEECH_CONVERSATION=1).

Safety

  • Hook always exits 0 on every path (verified: garbage stdin, empty transcript, no key, kill switch)
  • Without a recorder installed, barge-in silently degrades to normal playback
  • Offline unit tests (test_two_way.py): dB mapping sanity (quiet room ~100 dB-eq < 150 < shout ~180), floor gating on synthetic WAVs, handback JSON shape — all PASS

🤖 Generated with Codebuff

jeffrodrych7-stack and others added 2 commits September 28, 2026 11:12
While a response plays in background, the mic is monitored (sox/arecord/
ffmpeg). A sustained signal above SPEECH_BARGE_IN_DB (default 150,
SPL-equivalent - extreme by design so background noise can never interrupt)
kills playback and records the user's reply. ElevenLabs speech-to-text
(scribe_v1) transcribes it and the hook hands the words back to SirGent via
Stop-hook decision:block, so the conversation continues by voice.

New /voice command toggles the mode (flag-file based, wins over env).
Offline unit tests cover dB mapping, floor gating, and handback JSON.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <[email protected]>
@nastylex
nastylex merged commit e0294f2 into main Sep 28, 2026
@nastylex

Copy link
Copy Markdown
Owner

DoNE

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants