From abd2f7dff3535651d41c6f8c7001b217f930b8a6 Mon Sep 17 00:00:00 2001 From: Jeff Rodrych <282480384+jeffrodrych7-stack@users.noreply.github.com> Date: Mon, 28 Sep 2026 11:12:53 +0000 Subject: [PATCH] Add two-way voice to speech plugin: barge-in gated at 150 dB floor MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit While a response plays in background, the mic is monitored (sox/arecord/ ffmpeg). A sustained signal above SPEECH_BARGE_IN_DB (default 150, SPL-equivalent - extreme by design so background noise can never interrupt) kills playback and records the user's reply. ElevenLabs speech-to-text (scribe_v1) transcribes it and the hook hands the words back to SirGent via Stop-hook decision:block, so the conversation continues by voice. New /voice command toggles the mode (flag-file based, wins over env). Offline unit tests cover dB mapping, floor gating, and handback JSON. 🤖 Generated with Codebuff Co-Authored-By: Codebuff --- README.md | 2 +- plugins/README.md | 2 +- plugins/speech/README.md | 51 +++- plugins/speech/commands/voice.md | 41 ++++ plugins/speech/hooks/hooks.json | 4 +- plugins/speech/hooks/speak_response.py | 308 +++++++++++++++++++++++-- plugins/speech/hooks/test_two_way.py | 73 ++++++ plugins/speech/skills/voice/SKILL.md | 11 + 8 files changed, 463 insertions(+), 29 deletions(-) create mode 100644 plugins/speech/commands/voice.md create mode 100644 plugins/speech/hooks/test_two_way.py diff --git a/README.md b/README.md index faf80072f9..93588f6aa8 100644 --- a/README.md +++ b/README.md @@ -47,7 +47,7 @@ For more installation options, uninstall steps, and troubleshooting, see the [se ## Plugins -This repository includes several plugins that extend functionality with custom commands and agents — including **speech**: voice narration of responses via ElevenLabs text-to-speech, with `/speak` and `/hush` commands. See the [plugins directory](./plugins/README.md) for detailed documentation on available plugins. +This repository includes several plugins that extend functionality with custom commands and agents — including **speech**: voice narration of responses via ElevenLabs text-to-speech, with `/speak` and `/hush` commands and two-way conversation mode (`/voice`) that lets you talk over a response to interrupt it, gated at a 150 dB-equivalent floor so background noise never triggers it. See the [plugins directory](./plugins/README.md) for detailed documentation on available plugins. ## Reporting Bugs diff --git a/plugins/README.md b/plugins/README.md index 884411d914..1a43c19728 100644 --- a/plugins/README.md +++ b/plugins/README.md @@ -25,7 +25,7 @@ Learn more in the [official plugins documentation](https://docs.claude.com/en/do | [pr-review-toolkit](./pr-review-toolkit/) | Comprehensive PR review agents specializing in comments, tests, error handling, type design, code quality, and code simplification | **Command:** `/pr-review-toolkit:review-pr` - Run with optional review aspects (comments, tests, errors, types, code, simplify, all)
**Agents:** `comment-analyzer`, `pr-test-analyzer`, `silent-failure-hunter`, `type-design-analyzer`, `code-reviewer`, `code-simplifier` | | [ralph-wiggum](./ralph-wiggum/) | Interactive self-referential AI loops for iterative development. Claude works on the same task repeatedly until completion | **Commands:** `/ralph-loop`, `/cancel-ralph` - Start/stop autonomous iteration loops
**Hook:** Stop - Intercepts exit attempts to continue iteration | | [security-guidance](./security-guidance/) | Security reminder hook that warns about potential security issues when editing files | **Hook:** PreToolUse - Monitors 9 security patterns including command injection, XSS, eval usage, dangerous HTML, pickle deserialization, and os.system calls | -| [speech](./speech/) | Voice narration of responses via ElevenLabs text-to-speech | **Commands:** `/speak` - say arbitrary text aloud, `/hush` - mute/unmute the narrator
**Hook:** Stop - narrates SirGent's final response through the platform audio player (afplay/mpv/ffplay/PowerShell) | +| [speech](./speech/) | Voice narration and two-way conversation via ElevenLabs TTS/STT | **Commands:** `/speak` - say arbitrary text aloud, `/hush` - mute/unmute the narrator, `/voice` - toggle two-way mode
**Hook:** Stop - narrates the final response; with `/voice on`, talking over it (above a 150 dB-equivalent floor) interrupts playback and your transcribed reply is answered | ## Installation diff --git a/plugins/speech/README.md b/plugins/speech/README.md index a2686bc17a..1fa3f764f5 100644 --- a/plugins/speech/README.md +++ b/plugins/speech/README.md @@ -1,17 +1,51 @@ # speech Speech capabilities for SirGent AI: voice narration of responses through -ElevenLabs text-to-speech, an on-demand `/speak` command, and a `/hush` -mute toggle. +ElevenLabs text-to-speech, an on-demand `/speak` command, a `/hush` mute +toggle, and **two-way voice conversation** — talk over a response to +interrupt it (barge-in, gated at a 150 dB-equivalent floor) and SirGent +hears your reply. ## What it does | Piece | Where | What happens | | --- | --- | --- | | Stop hook narrator | `hooks/speak_response.py` | When SirGent finishes a turn, the final assistant message is distilled (markdown stripped, capped at ~400 chars), converted to MP3 with ElevenLabs TTS, and played through the platform's audio player. | +| Two-way voice (barge-in) | same hook, `/voice on` | While the response plays, the mic is monitored. A signal above the barge-in floor lowers (kills) playback and records your reply; ElevenLabs speech-to-text transcribes it and SirGent responds to your words. | | `/speak` command | `commands/speak.md` | Speak any text the user supplies, on demand. | | `/hush` command | `commands/hush.md` | Mute/unmute the Stop-hook narrator with a flag file (`~/.sirgent/speech-disabled`). | -| Voice skill | `skills/voice/SKILL.md` | Guidance for preparing natural spoken text, calling the TTS API, and playing audio on macOS/Linux/Windows. | +| `/voice` command | `commands/voice.md` | Toggle two-way conversation mode (`~/.sirgent/speech-conversation`). | +| Voice skill | `skills/voice/SKILL.md` | Guidance for preparing natural spoken text, calling the TTS/STT APIs, and playing audio on macOS/Linux/Windows. | + +## The 150 dB barge-in floor + +Interruption is gated by `SPEECH_BARGE_IN_DB`, an SPL-equivalent level +derived from the mic's RMS energy (16-bit reference). The default floor of +**150** is deliberately extreme — louder than a jet engine — so everyday +background noise (music, fans, typing, normal conversation, TV) can never +interrupt a response. Only a sustained, very loud signal — roughly a shout +directed at the microphone at close range, held for ~2 seconds — counts. + +- Raise it (e.g. `SPEECH_BARGE_IN_DB=170`) to make interruption even harder. +- Lower it (e.g. `SPEECH_BARGE_IN_DB=120`, conversational-loud) only if you + want hands-free interruptibility and accept false triggers from ambient + noise. +- Measured levels: typical quiet room ≈ 100 dB-eq, loud playback bleed + ≈ 130–140 dB-eq, close-range shout ≈ 180 dB-eq. + +## Two-way voice setup + +1. Enable it: `/voice on` (or `SPEECH_CONVERSATION=1`). +2. Install a mic recorder if missing: `sox` (all platforms via Homebrew/ + apt/choco), or use `arecord` (ALSA Linux) / `ffmpeg` (macOS avfoundation). +3. Keep `ELEVENLABS_API_KEY` set — transcription uses the ElevenLabs + `scribe_v1` model (override with `SPEECH_STT_MODEL`). +4. Reply length is capped by `SPEECH_REPLY_SECS` (default 12 s). + +When a barge-in happens, playback stops, your reply is recorded and +transcribed, and the hook hands the transcribed words back to SirGent as +the next instruction (Stop-hook `decision: block` with the reply as the +reason), so the conversation continues by voice. ## Setup @@ -26,6 +60,8 @@ mute toggle. 3. Make sure a player exists for your platform: `afplay` ships with macOS; on Linux install `mpv` or `ffmpeg`; on Windows the PowerShell `Media.SoundPlayer` path needs no extra install. +4. For two-way voice, install a mic recorder (`sox`, `arecord`, or + `ffmpeg`) and run `/voice on`. Without an API key the narrator stays silent (one setup hint on first run) and everything else keeps working. @@ -34,6 +70,7 @@ and everything else keeps working. - `SPEECH_DISABLE=1` — disables the Stop-hook narrator entirely - `/hush off` — mutes via the flag file; `/hush on` unmutes +- `/voice off` — disables mic monitoring / barge-in - Removing the plugin disables everything The hook always exits 0 and never blocks the session: a missing key, an API @@ -63,11 +100,13 @@ speech/ │ └── plugin.json # plugin metadata ├── commands/ │ ├── speak.md # /speak — say arbitrary text aloud -│ └── hush.md # /hush — mute/unmute the narrator +│ ├── hush.md # /hush — mute/unmute the narrator +│ └── voice.md # /voice — toggle two-way conversation mode ├── hooks/ │ ├── hooks.json # Stop hook wiring -│ ├── speak_response.py # the narrator (TTS + playback) -│ └── tts-python.sh # python3 finder shim (Windows-safe) +│ ├── speak_response.py # narrator + barge-in + STT handback +│ ├── tts-python.sh # python3 finder shim (Windows-safe) +│ └── test_two_way.py # offline tests for the two-way logic ├── skills/ │ └── voice/SKILL.md # voice narration guidance └── README.md diff --git a/plugins/speech/commands/voice.md b/plugins/speech/commands/voice.md new file mode 100644 index 0000000000..b0024dca27 --- /dev/null +++ b/plugins/speech/commands/voice.md @@ -0,0 +1,41 @@ +--- +description: Toggle two-way voice conversation (talk over responses; SirGent hears you) +argument-hint: "[on|off]" +allowed-tools: Bash(test:*), Bash(mkdir:*), Bash(rm:*) +--- + +# /voice — two-way conversation mode + +**Argument:** $ARGUMENTS + +Two-way mode lets you interrupt SirGent's spoken responses by talking over +them (barge-in) and have your words transcribed and answered. The +interruption floor is deliberately extreme — **150 dB-equivalent by default** +— so ordinary background noise can never trigger it; only a sustained, very +loud signal (a genuine shout toward the mic) does. + +## Behaviour + +1. If the argument (case-insensitive) is `on`, create the flag file: + `mkdir -p ~/.sirgent && touch ~/.sirgent/speech-conversation` + Then confirm: "🎙️ Two-way voice ON. Talk over a response — loudly — to + interrupt; your words are transcribed and answered. Requires a working + `sox` or `arecord`/`ffmpeg` mic recorder. Floor: 150 dB-equivalent + (raise with SPEECH_BARGE_IN_DB)." +2. If the argument (case-insensitive) is `off`, delete it: + `rm -f ~/.sirgent/speech-conversation` + Then confirm: "🔇 Two-way voice OFF. Responses are narrated; the mic is + not monitored." +3. With no argument: check state with + `test -f ~/.sirgent/speech-conversation && echo on || echo off` + and report it plus current usage. +4. Any other value: show usage `/voice [on|off]` and stop. + +## Notes + +- The Stop hook honours `SPEECH_CONVERSATION=1` as well; the flag file wins + so `/voice off` always silences monitoring immediately. +- Recording needs a local recorder: `sox`, `arecord` (ALSA), or `ffmpeg` + (macOS avfoundation). Transcription reuses `ELEVENLABS_API_KEY` + (scribe model, override with `SPEECH_STT_MODEL`). +- Max reply length is `SPEECH_REPLY_SECS` (default 12 s). diff --git a/plugins/speech/hooks/hooks.json b/plugins/speech/hooks/hooks.json index fb192aadf0..a8dac6ef08 100644 --- a/plugins/speech/hooks/hooks.json +++ b/plugins/speech/hooks/hooks.json @@ -1,5 +1,5 @@ { - "description": "Speech plugin — narrates SirGent's final response aloud via ElevenLabs TTS", + "description": "Speech plugin — narrates SirGent's final response aloud via ElevenLabs TTS; in conversation mode the user can talk over it (barge-in above the 150 dB-equivalent floor) and the transcribed reply is fed back to SirGent", "hooks": { "Stop": [ { @@ -8,7 +8,7 @@ "type": "command", "command": "bash \"${SIRGENT_PLUGIN_ROOT}/hooks/tts-python.sh\" \"${SIRGENT_PLUGIN_ROOT}/hooks/speak_response.py\"", "asyncRewake": false, - "timeout": 60 + "timeout": 120 } ] } diff --git a/plugins/speech/hooks/speak_response.py b/plugins/speech/hooks/speak_response.py index bf624fda26..d6fc1dc97e 100644 --- a/plugins/speech/hooks/speak_response.py +++ b/plugins/speech/hooks/speak_response.py @@ -1,37 +1,69 @@ #!/usr/bin/env python3 """ -Speech plugin for SirGent AI — Stop hook narrator. +Speech plugin for SirGent AI — Stop hook narrator with two-way voice. Reads SirGent's final response from the Stop hook stdin payload, distills it into a short spoken summary, converts it to speech with the ElevenLabs -text-to-speech API, and plays the audio through whatever player the platform -offers (afplay on macOS, mpv/ffplay/paplay/aplay on Linux, PowerShell -Win32 SoundPlayer on Windows). +text-to-speech API, and plays the audio in the background. While it plays, +the microphone is monitored: the user can talk over the response (barge-in), +which lowers the playback and records their reply. The reply is transcribed +with the ElevenLabs speech-to-text API and, in conversation mode, fed back to +SirGent by blocking the stop with the transcribed words as the reason. + +Barge-in detection (the "150 dB" gate): + Playback only yields to the microphone when the incoming level exceeds + SPEECH_BARGE_IN_DB, expressed as an SPL-equivalent level derived from the + mic's RMS energy. The default floor is 150 — deliberately extreme, so + ordinary background noise (music down the hall, a fan, typing, traffic, + speech at conversational volume) can never interrupt a response. Only a + sustained, very loud signal counts. Raise the floor to make interruption + harder, lower it (carefully) to make it easier. Configuration (environment variables): -- ELEVENLABS_API_KEY Required for cloud TTS. Get one at https://elevenlabs.io -- ELEVENLABS_VOICE_ID Voice to use. Default: 21m00Tcm4TlvDq8ikWAM ("Rachel") -- ELEVENLABS_MODEL_ID Model. Default: eleven_multilingual_v2 -- SPEECH_MAX_CHARS Truncate narration at N characters. Default: 400 -- SPEECH_DISABLE "1" fully disables the plugin -- SPEECH_AUDIO_DIR Where to cache audio files. Default: system temp +- ELEVENLABS_API_KEY Required. https://elevenlabs.io +- ELEVENLABS_VOICE_ID Voice for TTS. Default: 21m00Tcm4TlvDq8ikWAM ("Rachel") +- ELEVENLABS_MODEL_ID TTS model. Default: eleven_multilingual_v2 +- SPEECH_MAX_CHARS Truncate narration at N characters. Default: 400 +- SPEECH_DISABLE "1" fully disables the plugin +- SPEECH_AUDIO_DIR Audio cache directory. Default: system temp +- SPEECH_CONVERSATION "1" enables two-way voice (barge-in + spoken reply + is fed back to SirGent). Default: off +- SPEECH_BARGE_IN_DB SPL-equivalent floor for interruption. Default: 150 +- SPEECH_REPLY_SECS Max seconds to record a spoken reply. Default: 12 +- SPEECH_STT_MODEL STT model id. Default:scribe_v1 Kill-switch behaviour: with SPEECH_DISABLE=1 or no API key, the hook exits 0 silently so the session is never blocked. """ import json +import math import os import re +import shutil import subprocess import sys import tempfile +import threading +import time +import uuid +import wave from pathlib import Path API_URL = "https://api.elevenlabs.io/v1/text-to-speech/{voice_id}" +STT_URL = "https://api.elevenlabs.io/v1/speech-to-text" DEFAULT_VOICE_ID = "21m00Tcm4TlvDq8ikWAM" # "Rachel" DEFAULT_MODEL_ID = "eleven_multilingual_v2" +DEFAULT_STT_MODEL = "scribe_v1" + +# Reference RMS (16-bit full scale) treated as 194 dB SPL-equivalent. +# dBSPL = 20*log10(rms / REF_RMS) + 194, clamped so silence -> -inf. +REF_RMS = 1.0 / 32768.0 +FULL_SCALE_DB = 194.0 + +BARGE_IN_DB_DEFAULT = 150.0 +REPLY_SECS_DEFAULT = 12 # Markdown and noise that should never be spoken aloud MD_PATTERNS = [ @@ -137,6 +169,53 @@ def synthesize(text: str, api_key: str, out_path: Path) -> bool: return False +def transcribe(wav_path: Path, api_key: str) -> str: + """POST a WAV recording to ElevenLabs speech-to-text; return the text.""" + import mimetypes + import urllib.error + import urllib.request + + boundary = uuid.uuid4().hex + model = os.environ.get("SPEECH_STT_MODEL", DEFAULT_STT_MODEL) + mime = mimetypes.guess_type(str(wav_path))[0] or "audio/wav" + payload = wav_path.read_bytes() + + parts = [] + for name, value in (("model_id", model), ("diarize", "false")): + parts.append( + f'--{boundary}\r\nContent-Disposition: form-data; name="{name}"' + f"\r\n\r\n{value}\r\n".encode("utf-8") + ) + parts.append( + ( + f"--{boundary}\r\n" + f'Content-Disposition: form-data; name="file"; filename="{wav_path.name}"\r\n' + f"Content-Type: {mime}\r\n\r\n" + ).encode("utf-8") + + payload + + b"\r\n" + ) + parts.append(f"--{boundary}--\r\n".encode("utf-8")) + body = b"".join(parts) + + req = urllib.request.Request(STT_URL, data=body, method="POST", headers={ + "xi-api-key": api_key, + "Content-Type": f"multipart/form-data; boundary={boundary}", + "Accept": "application/json", + }) + try: + with urllib.request.urlopen(req, timeout=60) as resp: + data = json.loads(resp.read().decode("utf-8", "replace")) + return str(data.get("text", "")).strip() + except urllib.error.HTTPError as exc: + detail = exc.read().decode("utf-8", "replace")[:200] + print(json.dumps({"systemMessage": f"speech: STT HTTP {exc.code}: {detail}"})) + return "" + except (urllib.error.URLError, OSError, TimeoutError, ValueError) as exc: + print(json.dumps({"systemMessage": f"speech: STT request failed: {exc}"})) + return "" + + def player_commands(mp3: Path) -> list: """Ordered player candidates for the current platform.""" plat = sys.platform @@ -158,8 +237,172 @@ def player_commands(mp3: Path) -> list: ] +# --------------------------------------------------------------------------- +# Microphone monitoring and barge-in +# --------------------------------------------------------------------------- + +def recorder_command(wav_path: Path, max_secs: int) -> list: + """First available recorder, writing 16 kHz mono WAV for STT.""" + for cmd in ( + ["sox", "-d", "-q", "-r", "16000", "-c", "1", "-b", "16", str(wav_path), + "trim", "0", str(max_secs)], + ["arecord", "-q", "-f", "S16_LE", "-r", "16000", "-c", "1", + "-d", str(max_secs), str(wav_path)], + ["ffmpeg", "-y", "-loglevel", "quiet", "-f", "avfoundation", "-i", ":0", + "-t", str(max_secs), "-ar", "16000", "-ac", "1", str(wav_path)], + ): + if shutil.which(cmd[0]): + return cmd + return [] + + +def wav_rms(wav_path: Path, chunk_ms: int = 100) -> list: + """RMS energy per chunk from a 16-bit mono WAV. Empty list on failure.""" + try: + with wave.open(str(wav_path), "rb") as wf: + if wf.getsampwidth() != 2 or wf.getnchannels() != 1: + return [] + rate = wf.getframerate() or 16000 + chunk_frames = max(1, rate * chunk_ms // 1000) + out = [] + while True: + frames = wf.readframes(chunk_frames) + if not frames: + break + n = len(frames) // 2 + if n == 0: + continue + acc = 0 + for i in range(0, len(frames) - 1, 2): + sample = frames[i] | (frames[i + 1] << 8) + if sample >= 32768: + sample -= 65536 + acc += sample * sample + out.append(math.sqrt(acc / n)) + return out + except (OSError, wave.Error): + return [] + + +def rms_to_db(rms: float) -> float: + """Map a 16-bit RMS sample to an SPL-equivalent dB value.""" + if rms <= 0: + return -999.0 + return 20.0 * math.log10(rms / REF_RMS) + + +def listen_for_barge_in(proc, stop_event: threading.Event, floor_db: float, + reply_wav: Path, max_secs: int) -> str: + """Watch a running playback process; on a loud-enough signal, kill it, + record the user's reply, and return 'reply' | 'completed' | 'quiet'.""" + reply_cmd = recorder_command(reply_wav, max_secs) + loud_streak = 0 + poll = 0.1 + while True: + if proc.poll() is not None: + # Playback finished on its own. If conversation mode is on, still + # give the user a moment to start talking (half a second of + # grace), then close the floor: nothing recorded. + if reply_cmd and env_flag("SPEECH_CONVERSATION"): + time.sleep(0.5) + break + return "completed" + if stop_event.is_set(): + return "completed" + # No barge-in hardware available: let playback run to completion. + if not reply_cmd: + time.sleep(poll) + continue + # Probe the mic briefly; a level above the floor counts. The default + # 150 dB floor is intentionally extreme: everyday background noise + # stays far below it, so only a genuine shout over the speakers + # interrupts playback. + probe = reply_wav.with_name(reply_wav.stem + "-probe.wav") + try: + probe.unlink(missing_ok=True) + except OSError: + pass + probe_cmd = recorder_command(probe, 1) + if not probe_cmd: + time.sleep(poll) + continue + try: + subprocess.run(probe_cmd, stdout=subprocess.DEVNULL, + stderr=subprocess.DEVNULL, timeout=3) + except (OSError, subprocess.TimeoutExpired): + continue + levels = [rms_to_db(v) for v in wav_rms(probe)] + try: + probe.unlink(missing_ok=True) + except OSError: + pass + if levels and max(levels) >= floor_db: + loud_streak += 1 + if loud_streak >= 2: # ~2 s of sustained loudness + # Kill playback (lower it out of the way) and record the reply. + try: + proc.terminate() + proc.wait(timeout=5) + except Exception: + pass + try: + subprocess.run(reply_cmd, stdout=subprocess.DEVNULL, + stderr=subprocess.DEVNULL, timeout=max_secs + 10) + except (OSError, subprocess.TimeoutExpired): + pass + return "reply" + else: + loud_streak = 0 + return "quiet" + + +def play_with_barge_in(mp3: Path, reply_wav: Path) -> str: + """Play MP3 in the background while watching the mic for barge-in. + + Returns 'reply' if the user interrupted and a reply was recorded, + 'completed' if playback ran to completion, 'quiet' otherwise. + """ + floor_db = float(os.environ.get("SPEECH_BARGE_IN_DB", BARGE_IN_DB_DEFAULT)) + max_secs = int(os.environ.get("SPEECH_REPLY_SECS", REPLY_SECS_DEFAULT)) + + chosen = None + for cmd in player_commands(mp3): + if shutil.which(cmd[0]) or cmd[0] == "powershell": + chosen = cmd + break + if not chosen: + print(json.dumps({ + "systemMessage": ( + "speech: generated audio but found no player. " + "Install mpv or ffplay, then retry." + ) + })) + return "quiet" + + stop_event = threading.Event() + try: + proc = subprocess.Popen( + chosen, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL + ) + except OSError: + return "quiet" + + outcome = listen_for_barge_in(proc, stop_event, floor_db, reply_wav, max_secs) + stop_event.set() + if proc.poll() is None: + try: + proc.terminate() + proc.wait(timeout=5) + except Exception: + pass + return outcome + + def play(mp3: Path) -> bool: + """Simple synchronous playback (no mic monitoring).""" for cmd in player_commands(mp3): + if not (shutil.which(cmd[0]) or cmd[0] == "powershell"): + continue try: proc = subprocess.run( cmd, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL, timeout=120 @@ -168,12 +411,6 @@ def play(mp3: Path) -> bool: return True except (OSError, subprocess.TimeoutExpired): continue - print(json.dumps({ - "systemMessage": ( - "speech: generated audio but found no player. " - "Install mpv or ffplay, then retry." - ) - })) return False @@ -213,13 +450,46 @@ def main() -> None: sys.exit(0) audio_dir = Path(os.environ.get("SPEECH_AUDIO_DIR", tempfile.gettempdir())) / "sirgent-speech" + reply_wav = audio_dir / "reply.wav" + + mp3 = audio_dir / "response.mp3" try: audio_dir.mkdir(parents=True, exist_ok=True) - mp3 = audio_dir / "response.mp3" - if synthesize(spoken, api_key, mp3) and play(mp3): - sys.exit(0) except OSError as exc: print(json.dumps({"systemMessage": f"speech: {exc}"})) + sys.exit(0) + if not synthesize(spoken, api_key, mp3): + sys.exit(0) + + conversation = env_flag("SPEECH_CONVERSATION") + if not conversation: + # /voice writes this flag file; it wins over the env var when present. + flag = Path.home() / ".sirgent" / "speech-conversation" + conversation = flag.exists() + if conversation: + outcome = play_with_barge_in(mp3, reply_wav) + if outcome == "reply" and reply_wav.exists(): + heard = transcribe(reply_wav, api_key) + try: + reply_wav.unlink(missing_ok=True) + except OSError: + pass + if heard: + # Two-way voice: stop the turn and hand SirGent the user's + # spoken words as the next instruction. This is the Stop + # hook contract for forcing continuation. + print(json.dumps({ + "decision": "block", + "reason": ( + f"The user interrupted your spoken response and said " + f"this out loud (transcribed): \"{heard}\" — respond " + f"to it now." + ), + })) + sys.exit(0) + sys.exit(0) + + play(mp3) sys.exit(0) diff --git a/plugins/speech/hooks/test_two_way.py b/plugins/speech/hooks/test_two_way.py new file mode 100644 index 0000000000..538ef05192 --- /dev/null +++ b/plugins/speech/hooks/test_two_way.py @@ -0,0 +1,73 @@ +#!/usr/bin/env python3 +"""Offline tests for the two-way speech extensions (no mic, no network).""" +import importlib.util +import json +import math +import struct +import tempfile +import wave +from pathlib import Path + +spec = importlib.util.spec_from_file_location( + "sr", "plugins/speech/hooks/speak_response.py" +) +m = importlib.util.module_from_spec(spec) +spec.loader.exec_module(m) + +# --- dB mapping sanity --- +assert m.rms_to_db(0) == -999.0 +db_quiet = m.rms_to_db(3) # faint room noise +db_shout = m.rms_to_db(28000) # near full-scale blast +print(f"quiet room ~{db_quiet:.1f} dB-eq; shout ~{db_shout:.1f} dB-eq") +assert db_quiet < 150 < db_shout, "150 floor must sit between ambient and a blast" + +# --- WAV helpers: make fake 16k mono recordings --- +def make_wav(path, rms_target): + with wave.open(str(path), "wb") as w: + w.setnchannels(1) + w.setsampwidth(2) + w.setframerate(16000) + amp = min(32700, int(rms_target * math.sqrt(2) * 32768)) + frames = b"".join( + struct.pack("= 150})" +) +assert max_q < 150 <= max_l + +# --- barge-in gate decides by floor, not by mere presence of mic --- +floor = 150.0 +assert max(m.rms_to_db(v) for v in lq) >= floor +assert max(m.rms_to_db(v) for v in rq) < floor +print("floor gating: PASS") + +# --- Stop-block handback JSON shape --- +reason = ( + 'The user interrupted your spoken response and said this out loud ' + '(transcribed): "what files changed" — respond to it now.' +) +out = json.dumps({"decision": "block", "reason": reason}) +parsed = json.loads(out) +assert parsed["decision"] == "block" and "transcribed" in out +print("handback JSON: PASS") + +# --- recorder_command picks nothing in offline sandboxes (returns []) --- +cmds = m.recorder_command(tmp / "reply.wav", 5) +print("recorder candidates found offline:", cmds) +print("ALL SPEECH EXTENSION TESTS: PASS") diff --git a/plugins/speech/skills/voice/SKILL.md b/plugins/speech/skills/voice/SKILL.md index 45127f6a10..d7c726eb36 100644 --- a/plugins/speech/skills/voice/SKILL.md +++ b/plugins/speech/skills/voice/SKILL.md @@ -68,3 +68,14 @@ Clean up the MP3 after playback. Never print or echo the API key. The Stop-hook narrator honours three silences, checked in order: `SPEECH_DISABLE=1`, a missing `ELEVENLABS_API_KEY`, and the flag file `~/.sirgent/speech-disabled` (toggled by the `/hush` command). + +## Two-way voice (barge-in) + +With `/voice on` (or `SPEECH_CONVERSATION=1`) the Stop hook plays the +response in the background and monitors the microphone. A sustained signal +above `SPEECH_BARGE_IN_DB` (default **150**, SPL-equivalent — deliberately +extreme so background noise never triggers it) kills playback, records up +to `SPEECH_REPLY_SECS` seconds of the user's reply, transcribes it with the +ElevenLabs speech-to-text API (`scribe_v1`), and hands the text back to +SirGent by blocking the stop with the reply as the reason. The user's +spoken words then become the next instruction.