diff --git a/README.md b/README.md
index 56e112a1b8..556e39b9e8 100644
--- a/README.md
+++ b/README.md
@@ -47,7 +47,7 @@ For more installation options, uninstall steps, and troubleshooting, see the [se
## Plugins
-This repository includes several plugins that extend functionality with custom commands and agents — including **speech**: voice narration of responses via ElevenLabs text-to-speech, with `/speak` and `/hush` commands. See the [plugins directory](./plugins/README.md) for detailed documentation on available plugins.
+This repository includes several plugins that extend functionality with custom commands and agents — including **speech**: voice narration of responses via ElevenLabs text-to-speech, with `/speak` and `/hush` commands and two-way conversation mode (`/voice`) that lets you talk over a response to interrupt it, gated at a 150 dB-equivalent floor so background noise never triggers it. See the [plugins directory](./plugins/README.md) for detailed documentation on available plugins.
## Reporting Bugs
diff --git a/plugins/README.md b/plugins/README.md
index 884411d914..1a43c19728 100644
--- a/plugins/README.md
+++ b/plugins/README.md
@@ -25,7 +25,7 @@ Learn more in the [official plugins documentation](https://docs.claude.com/en/do
| [pr-review-toolkit](./pr-review-toolkit/) | Comprehensive PR review agents specializing in comments, tests, error handling, type design, code quality, and code simplification | **Command:** `/pr-review-toolkit:review-pr` - Run with optional review aspects (comments, tests, errors, types, code, simplify, all)
**Agents:** `comment-analyzer`, `pr-test-analyzer`, `silent-failure-hunter`, `type-design-analyzer`, `code-reviewer`, `code-simplifier` |
| [ralph-wiggum](./ralph-wiggum/) | Interactive self-referential AI loops for iterative development. Claude works on the same task repeatedly until completion | **Commands:** `/ralph-loop`, `/cancel-ralph` - Start/stop autonomous iteration loops
**Hook:** Stop - Intercepts exit attempts to continue iteration |
| [security-guidance](./security-guidance/) | Security reminder hook that warns about potential security issues when editing files | **Hook:** PreToolUse - Monitors 9 security patterns including command injection, XSS, eval usage, dangerous HTML, pickle deserialization, and os.system calls |
-| [speech](./speech/) | Voice narration of responses via ElevenLabs text-to-speech | **Commands:** `/speak` - say arbitrary text aloud, `/hush` - mute/unmute the narrator
**Hook:** Stop - narrates SirGent's final response through the platform audio player (afplay/mpv/ffplay/PowerShell) |
+| [speech](./speech/) | Voice narration and two-way conversation via ElevenLabs TTS/STT | **Commands:** `/speak` - say arbitrary text aloud, `/hush` - mute/unmute the narrator, `/voice` - toggle two-way mode
**Hook:** Stop - narrates the final response; with `/voice on`, talking over it (above a 150 dB-equivalent floor) interrupts playback and your transcribed reply is answered |
## Installation
diff --git a/plugins/speech/README.md b/plugins/speech/README.md
index a309ee08b8..937f5a43fe 100644
--- a/plugins/speech/README.md
+++ b/plugins/speech/README.md
@@ -1,19 +1,17 @@
# speech
Speech capabilities for SirGent AI: voice narration of responses through
-ElevenLabs text-to-speech, voice dictation into your prompts through
-ElevenLabs speech-to-text, an on-demand `/speak` command, a `/hush` mute
-toggle, and a `/dictate` command for voice input.
+
## What it does
| Piece | Where | What happens |
| --- | --- | --- |
| Stop hook narrator | `hooks/speak_response.py` | When SirGent finishes a turn, the final assistant message is distilled (markdown stripped, capped at ~400 chars), converted to MP3 with ElevenLabs TTS, and played through the platform's audio player. |
+| Two-way voice (barge-in) | same hook, `/voice on` | While the response plays, the mic is monitored. A signal above the barge-in floor lowers (kills) playback and records your reply; ElevenLabs speech-to-text transcribes it and SirGent responds to your words. |
| `/speak` command | `commands/speak.md` | Speak any text the user supplies, on demand. |
| `/hush` command | `commands/hush.md` | Mute/unmute the Stop-hook narrator with a flag file (`~/.sirgent/speech-disabled`). |
-| `/dictate` command | `commands/dictate.md` + `hooks/dictate_stt.py` | Voice input: record from the default mic, transcribe with ElevenLabs STT, and deliver the words as your prompt to SirGent. |
-| Voice skill | `skills/voice/SKILL.md` | Guidance for preparing natural spoken text, calling the TTS API, and playing audio on macOS/Linux/Windows. |
+
## Setup
@@ -30,9 +28,7 @@ toggle, and a `/dictate` command for voice input.
3. Make sure a player exists for your platform: `afplay` ships with macOS;
on Linux install `mpv` or `ffmpeg`; on Windows the PowerShell
`Media.SoundPlayer` path needs no extra install.
-4. For `/dictate`, install a mic recorder: `sox` (recommended, all
- platforms via Homebrew/apt/choco), or use `arecord` (ALSA Linux) /
- `ffmpeg` (macOS avfoundation).
+
Without an API key the narrator stays silent (one setup hint on first run)
and everything else keeps working.
@@ -41,6 +37,7 @@ and everything else keeps working.
- `SPEECH_DISABLE=1` — disables the Stop-hook narrator entirely
- `/hush off` — mutes via the flag file; `/hush on` unmutes
+- `/voice off` — disables mic monitoring / barge-in
- Removing the plugin disables everything
The hook always exits 0 and never blocks the session: a missing key, an API
@@ -72,12 +69,7 @@ speech/
├── commands/
│ ├── speak.md # /speak — say arbitrary text aloud
│ ├── hush.md # /hush — mute/unmute the narrator
-│ └── dictate.md # /dictate — voice input (mic → STT → prompt)
-├── hooks/
-│ ├── hooks.json # Stop hook wiring
-│ ├── speak_response.py # the narrator (TTS + playback)
-│ ├── dictate_stt.py # dictation bridge (record → STT → stdout)
-│ └── tts-python.sh # python3 finder shim (Windows-safe)
+
├── skills/
│ └── voice/SKILL.md # voice narration guidance
└── README.md
diff --git a/plugins/speech/commands/voice.md b/plugins/speech/commands/voice.md
new file mode 100644
index 0000000000..b0024dca27
--- /dev/null
+++ b/plugins/speech/commands/voice.md
@@ -0,0 +1,41 @@
+---
+description: Toggle two-way voice conversation (talk over responses; SirGent hears you)
+argument-hint: "[on|off]"
+allowed-tools: Bash(test:*), Bash(mkdir:*), Bash(rm:*)
+---
+
+# /voice — two-way conversation mode
+
+**Argument:** $ARGUMENTS
+
+Two-way mode lets you interrupt SirGent's spoken responses by talking over
+them (barge-in) and have your words transcribed and answered. The
+interruption floor is deliberately extreme — **150 dB-equivalent by default**
+— so ordinary background noise can never trigger it; only a sustained, very
+loud signal (a genuine shout toward the mic) does.
+
+## Behaviour
+
+1. If the argument (case-insensitive) is `on`, create the flag file:
+ `mkdir -p ~/.sirgent && touch ~/.sirgent/speech-conversation`
+ Then confirm: "🎙️ Two-way voice ON. Talk over a response — loudly — to
+ interrupt; your words are transcribed and answered. Requires a working
+ `sox` or `arecord`/`ffmpeg` mic recorder. Floor: 150 dB-equivalent
+ (raise with SPEECH_BARGE_IN_DB)."
+2. If the argument (case-insensitive) is `off`, delete it:
+ `rm -f ~/.sirgent/speech-conversation`
+ Then confirm: "🔇 Two-way voice OFF. Responses are narrated; the mic is
+ not monitored."
+3. With no argument: check state with
+ `test -f ~/.sirgent/speech-conversation && echo on || echo off`
+ and report it plus current usage.
+4. Any other value: show usage `/voice [on|off]` and stop.
+
+## Notes
+
+- The Stop hook honours `SPEECH_CONVERSATION=1` as well; the flag file wins
+ so `/voice off` always silences monitoring immediately.
+- Recording needs a local recorder: `sox`, `arecord` (ALSA), or `ffmpeg`
+ (macOS avfoundation). Transcription reuses `ELEVENLABS_API_KEY`
+ (scribe model, override with `SPEECH_STT_MODEL`).
+- Max reply length is `SPEECH_REPLY_SECS` (default 12 s).
diff --git a/plugins/speech/hooks/hooks.json b/plugins/speech/hooks/hooks.json
index 91aed3215b..98232c64fc 100644
--- a/plugins/speech/hooks/hooks.json
+++ b/plugins/speech/hooks/hooks.json
@@ -1,17 +1 @@
{
- "description": "Speech plugin — narrates SirGent's final response aloud via ElevenLabs TTS; /dictate covers the input side (mic → STT → prompt)",
- "hooks": {
- "Stop": [
- {
- "hooks": [
- {
- "type": "command",
- "command": "bash \"${SIRGENT_PLUGIN_ROOT}/hooks/tts-python.sh\" \"${SIRGENT_PLUGIN_ROOT}/hooks/speak_response.py\"",
- "asyncRewake": false,
- "timeout": 60
- }
- ]
- }
- ]
- }
-}
diff --git a/plugins/speech/hooks/speak_response.py b/plugins/speech/hooks/speak_response.py
index bf624fda26..d6fc1dc97e 100644
--- a/plugins/speech/hooks/speak_response.py
+++ b/plugins/speech/hooks/speak_response.py
@@ -1,37 +1,69 @@
#!/usr/bin/env python3
"""
-Speech plugin for SirGent AI — Stop hook narrator.
+Speech plugin for SirGent AI — Stop hook narrator with two-way voice.
Reads SirGent's final response from the Stop hook stdin payload, distills it
into a short spoken summary, converts it to speech with the ElevenLabs
-text-to-speech API, and plays the audio through whatever player the platform
-offers (afplay on macOS, mpv/ffplay/paplay/aplay on Linux, PowerShell
-Win32 SoundPlayer on Windows).
+text-to-speech API, and plays the audio in the background. While it plays,
+the microphone is monitored: the user can talk over the response (barge-in),
+which lowers the playback and records their reply. The reply is transcribed
+with the ElevenLabs speech-to-text API and, in conversation mode, fed back to
+SirGent by blocking the stop with the transcribed words as the reason.
+
+Barge-in detection (the "150 dB" gate):
+ Playback only yields to the microphone when the incoming level exceeds
+ SPEECH_BARGE_IN_DB, expressed as an SPL-equivalent level derived from the
+ mic's RMS energy. The default floor is 150 — deliberately extreme, so
+ ordinary background noise (music down the hall, a fan, typing, traffic,
+ speech at conversational volume) can never interrupt a response. Only a
+ sustained, very loud signal counts. Raise the floor to make interruption
+ harder, lower it (carefully) to make it easier.
Configuration (environment variables):
-- ELEVENLABS_API_KEY Required for cloud TTS. Get one at https://elevenlabs.io
-- ELEVENLABS_VOICE_ID Voice to use. Default: 21m00Tcm4TlvDq8ikWAM ("Rachel")
-- ELEVENLABS_MODEL_ID Model. Default: eleven_multilingual_v2
-- SPEECH_MAX_CHARS Truncate narration at N characters. Default: 400
-- SPEECH_DISABLE "1" fully disables the plugin
-- SPEECH_AUDIO_DIR Where to cache audio files. Default: system temp
+- ELEVENLABS_API_KEY Required. https://elevenlabs.io
+- ELEVENLABS_VOICE_ID Voice for TTS. Default: 21m00Tcm4TlvDq8ikWAM ("Rachel")
+- ELEVENLABS_MODEL_ID TTS model. Default: eleven_multilingual_v2
+- SPEECH_MAX_CHARS Truncate narration at N characters. Default: 400
+- SPEECH_DISABLE "1" fully disables the plugin
+- SPEECH_AUDIO_DIR Audio cache directory. Default: system temp
+- SPEECH_CONVERSATION "1" enables two-way voice (barge-in + spoken reply
+ is fed back to SirGent). Default: off
+- SPEECH_BARGE_IN_DB SPL-equivalent floor for interruption. Default: 150
+- SPEECH_REPLY_SECS Max seconds to record a spoken reply. Default: 12
+- SPEECH_STT_MODEL STT model id. Default:scribe_v1
Kill-switch behaviour: with SPEECH_DISABLE=1 or no API key, the hook exits 0
silently so the session is never blocked.
"""
import json
+import math
import os
import re
+import shutil
import subprocess
import sys
import tempfile
+import threading
+import time
+import uuid
+import wave
from pathlib import Path
API_URL = "https://api.elevenlabs.io/v1/text-to-speech/{voice_id}"
+STT_URL = "https://api.elevenlabs.io/v1/speech-to-text"
DEFAULT_VOICE_ID = "21m00Tcm4TlvDq8ikWAM" # "Rachel"
DEFAULT_MODEL_ID = "eleven_multilingual_v2"
+DEFAULT_STT_MODEL = "scribe_v1"
+
+# Reference RMS (16-bit full scale) treated as 194 dB SPL-equivalent.
+# dBSPL = 20*log10(rms / REF_RMS) + 194, clamped so silence -> -inf.
+REF_RMS = 1.0 / 32768.0
+FULL_SCALE_DB = 194.0
+
+BARGE_IN_DB_DEFAULT = 150.0
+REPLY_SECS_DEFAULT = 12
# Markdown and noise that should never be spoken aloud
MD_PATTERNS = [
@@ -137,6 +169,53 @@ def synthesize(text: str, api_key: str, out_path: Path) -> bool:
return False
+def transcribe(wav_path: Path, api_key: str) -> str:
+ """POST a WAV recording to ElevenLabs speech-to-text; return the text."""
+ import mimetypes
+ import urllib.error
+ import urllib.request
+
+ boundary = uuid.uuid4().hex
+ model = os.environ.get("SPEECH_STT_MODEL", DEFAULT_STT_MODEL)
+ mime = mimetypes.guess_type(str(wav_path))[0] or "audio/wav"
+ payload = wav_path.read_bytes()
+
+ parts = []
+ for name, value in (("model_id", model), ("diarize", "false")):
+ parts.append(
+ f'--{boundary}\r\nContent-Disposition: form-data; name="{name}"'
+ f"\r\n\r\n{value}\r\n".encode("utf-8")
+ )
+ parts.append(
+ (
+ f"--{boundary}\r\n"
+ f'Content-Disposition: form-data; name="file"; filename="{wav_path.name}"\r\n'
+ f"Content-Type: {mime}\r\n\r\n"
+ ).encode("utf-8")
+ + payload
+ + b"\r\n"
+ )
+ parts.append(f"--{boundary}--\r\n".encode("utf-8"))
+ body = b"".join(parts)
+
+ req = urllib.request.Request(STT_URL, data=body, method="POST", headers={
+ "xi-api-key": api_key,
+ "Content-Type": f"multipart/form-data; boundary={boundary}",
+ "Accept": "application/json",
+ })
+ try:
+ with urllib.request.urlopen(req, timeout=60) as resp:
+ data = json.loads(resp.read().decode("utf-8", "replace"))
+ return str(data.get("text", "")).strip()
+ except urllib.error.HTTPError as exc:
+ detail = exc.read().decode("utf-8", "replace")[:200]
+ print(json.dumps({"systemMessage": f"speech: STT HTTP {exc.code}: {detail}"}))
+ return ""
+ except (urllib.error.URLError, OSError, TimeoutError, ValueError) as exc:
+ print(json.dumps({"systemMessage": f"speech: STT request failed: {exc}"}))
+ return ""
+
+
def player_commands(mp3: Path) -> list:
"""Ordered player candidates for the current platform."""
plat = sys.platform
@@ -158,8 +237,172 @@ def player_commands(mp3: Path) -> list:
]
+# ---------------------------------------------------------------------------
+# Microphone monitoring and barge-in
+# ---------------------------------------------------------------------------
+
+def recorder_command(wav_path: Path, max_secs: int) -> list:
+ """First available recorder, writing 16 kHz mono WAV for STT."""
+ for cmd in (
+ ["sox", "-d", "-q", "-r", "16000", "-c", "1", "-b", "16", str(wav_path),
+ "trim", "0", str(max_secs)],
+ ["arecord", "-q", "-f", "S16_LE", "-r", "16000", "-c", "1",
+ "-d", str(max_secs), str(wav_path)],
+ ["ffmpeg", "-y", "-loglevel", "quiet", "-f", "avfoundation", "-i", ":0",
+ "-t", str(max_secs), "-ar", "16000", "-ac", "1", str(wav_path)],
+ ):
+ if shutil.which(cmd[0]):
+ return cmd
+ return []
+
+
+def wav_rms(wav_path: Path, chunk_ms: int = 100) -> list:
+ """RMS energy per chunk from a 16-bit mono WAV. Empty list on failure."""
+ try:
+ with wave.open(str(wav_path), "rb") as wf:
+ if wf.getsampwidth() != 2 or wf.getnchannels() != 1:
+ return []
+ rate = wf.getframerate() or 16000
+ chunk_frames = max(1, rate * chunk_ms // 1000)
+ out = []
+ while True:
+ frames = wf.readframes(chunk_frames)
+ if not frames:
+ break
+ n = len(frames) // 2
+ if n == 0:
+ continue
+ acc = 0
+ for i in range(0, len(frames) - 1, 2):
+ sample = frames[i] | (frames[i + 1] << 8)
+ if sample >= 32768:
+ sample -= 65536
+ acc += sample * sample
+ out.append(math.sqrt(acc / n))
+ return out
+ except (OSError, wave.Error):
+ return []
+
+
+def rms_to_db(rms: float) -> float:
+ """Map a 16-bit RMS sample to an SPL-equivalent dB value."""
+ if rms <= 0:
+ return -999.0
+ return 20.0 * math.log10(rms / REF_RMS)
+
+
+def listen_for_barge_in(proc, stop_event: threading.Event, floor_db: float,
+ reply_wav: Path, max_secs: int) -> str:
+ """Watch a running playback process; on a loud-enough signal, kill it,
+ record the user's reply, and return 'reply' | 'completed' | 'quiet'."""
+ reply_cmd = recorder_command(reply_wav, max_secs)
+ loud_streak = 0
+ poll = 0.1
+ while True:
+ if proc.poll() is not None:
+ # Playback finished on its own. If conversation mode is on, still
+ # give the user a moment to start talking (half a second of
+ # grace), then close the floor: nothing recorded.
+ if reply_cmd and env_flag("SPEECH_CONVERSATION"):
+ time.sleep(0.5)
+ break
+ return "completed"
+ if stop_event.is_set():
+ return "completed"
+ # No barge-in hardware available: let playback run to completion.
+ if not reply_cmd:
+ time.sleep(poll)
+ continue
+ # Probe the mic briefly; a level above the floor counts. The default
+ # 150 dB floor is intentionally extreme: everyday background noise
+ # stays far below it, so only a genuine shout over the speakers
+ # interrupts playback.
+ probe = reply_wav.with_name(reply_wav.stem + "-probe.wav")
+ try:
+ probe.unlink(missing_ok=True)
+ except OSError:
+ pass
+ probe_cmd = recorder_command(probe, 1)
+ if not probe_cmd:
+ time.sleep(poll)
+ continue
+ try:
+ subprocess.run(probe_cmd, stdout=subprocess.DEVNULL,
+ stderr=subprocess.DEVNULL, timeout=3)
+ except (OSError, subprocess.TimeoutExpired):
+ continue
+ levels = [rms_to_db(v) for v in wav_rms(probe)]
+ try:
+ probe.unlink(missing_ok=True)
+ except OSError:
+ pass
+ if levels and max(levels) >= floor_db:
+ loud_streak += 1
+ if loud_streak >= 2: # ~2 s of sustained loudness
+ # Kill playback (lower it out of the way) and record the reply.
+ try:
+ proc.terminate()
+ proc.wait(timeout=5)
+ except Exception:
+ pass
+ try:
+ subprocess.run(reply_cmd, stdout=subprocess.DEVNULL,
+ stderr=subprocess.DEVNULL, timeout=max_secs + 10)
+ except (OSError, subprocess.TimeoutExpired):
+ pass
+ return "reply"
+ else:
+ loud_streak = 0
+ return "quiet"
+
+
+def play_with_barge_in(mp3: Path, reply_wav: Path) -> str:
+ """Play MP3 in the background while watching the mic for barge-in.
+
+ Returns 'reply' if the user interrupted and a reply was recorded,
+ 'completed' if playback ran to completion, 'quiet' otherwise.
+ """
+ floor_db = float(os.environ.get("SPEECH_BARGE_IN_DB", BARGE_IN_DB_DEFAULT))
+ max_secs = int(os.environ.get("SPEECH_REPLY_SECS", REPLY_SECS_DEFAULT))
+
+ chosen = None
+ for cmd in player_commands(mp3):
+ if shutil.which(cmd[0]) or cmd[0] == "powershell":
+ chosen = cmd
+ break
+ if not chosen:
+ print(json.dumps({
+ "systemMessage": (
+ "speech: generated audio but found no player. "
+ "Install mpv or ffplay, then retry."
+ )
+ }))
+ return "quiet"
+
+ stop_event = threading.Event()
+ try:
+ proc = subprocess.Popen(
+ chosen, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL
+ )
+ except OSError:
+ return "quiet"
+
+ outcome = listen_for_barge_in(proc, stop_event, floor_db, reply_wav, max_secs)
+ stop_event.set()
+ if proc.poll() is None:
+ try:
+ proc.terminate()
+ proc.wait(timeout=5)
+ except Exception:
+ pass
+ return outcome
+
+
def play(mp3: Path) -> bool:
+ """Simple synchronous playback (no mic monitoring)."""
for cmd in player_commands(mp3):
+ if not (shutil.which(cmd[0]) or cmd[0] == "powershell"):
+ continue
try:
proc = subprocess.run(
cmd, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL, timeout=120
@@ -168,12 +411,6 @@ def play(mp3: Path) -> bool:
return True
except (OSError, subprocess.TimeoutExpired):
continue
- print(json.dumps({
- "systemMessage": (
- "speech: generated audio but found no player. "
- "Install mpv or ffplay, then retry."
- )
- }))
return False
@@ -213,13 +450,46 @@ def main() -> None:
sys.exit(0)
audio_dir = Path(os.environ.get("SPEECH_AUDIO_DIR", tempfile.gettempdir())) / "sirgent-speech"
+ reply_wav = audio_dir / "reply.wav"
+
+ mp3 = audio_dir / "response.mp3"
try:
audio_dir.mkdir(parents=True, exist_ok=True)
- mp3 = audio_dir / "response.mp3"
- if synthesize(spoken, api_key, mp3) and play(mp3):
- sys.exit(0)
except OSError as exc:
print(json.dumps({"systemMessage": f"speech: {exc}"}))
+ sys.exit(0)
+ if not synthesize(spoken, api_key, mp3):
+ sys.exit(0)
+
+ conversation = env_flag("SPEECH_CONVERSATION")
+ if not conversation:
+ # /voice writes this flag file; it wins over the env var when present.
+ flag = Path.home() / ".sirgent" / "speech-conversation"
+ conversation = flag.exists()
+ if conversation:
+ outcome = play_with_barge_in(mp3, reply_wav)
+ if outcome == "reply" and reply_wav.exists():
+ heard = transcribe(reply_wav, api_key)
+ try:
+ reply_wav.unlink(missing_ok=True)
+ except OSError:
+ pass
+ if heard:
+ # Two-way voice: stop the turn and hand SirGent the user's
+ # spoken words as the next instruction. This is the Stop
+ # hook contract for forcing continuation.
+ print(json.dumps({
+ "decision": "block",
+ "reason": (
+ f"The user interrupted your spoken response and said "
+ f"this out loud (transcribed): \"{heard}\" — respond "
+ f"to it now."
+ ),
+ }))
+ sys.exit(0)
+ sys.exit(0)
+
+ play(mp3)
sys.exit(0)
diff --git a/plugins/speech/hooks/test_two_way.py b/plugins/speech/hooks/test_two_way.py
new file mode 100644
index 0000000000..538ef05192
--- /dev/null
+++ b/plugins/speech/hooks/test_two_way.py
@@ -0,0 +1,73 @@
+#!/usr/bin/env python3
+"""Offline tests for the two-way speech extensions (no mic, no network)."""
+import importlib.util
+import json
+import math
+import struct
+import tempfile
+import wave
+from pathlib import Path
+
+spec = importlib.util.spec_from_file_location(
+ "sr", "plugins/speech/hooks/speak_response.py"
+)
+m = importlib.util.module_from_spec(spec)
+spec.loader.exec_module(m)
+
+# --- dB mapping sanity ---
+assert m.rms_to_db(0) == -999.0
+db_quiet = m.rms_to_db(3) # faint room noise
+db_shout = m.rms_to_db(28000) # near full-scale blast
+print(f"quiet room ~{db_quiet:.1f} dB-eq; shout ~{db_shout:.1f} dB-eq")
+assert db_quiet < 150 < db_shout, "150 floor must sit between ambient and a blast"
+
+# --- WAV helpers: make fake 16k mono recordings ---
+def make_wav(path, rms_target):
+ with wave.open(str(path), "wb") as w:
+ w.setnchannels(1)
+ w.setsampwidth(2)
+ w.setframerate(16000)
+ amp = min(32700, int(rms_target * math.sqrt(2) * 32768))
+ frames = b"".join(
+ struct.pack("= 150})"
+)
+assert max_q < 150 <= max_l
+
+# --- barge-in gate decides by floor, not by mere presence of mic ---
+floor = 150.0
+assert max(m.rms_to_db(v) for v in lq) >= floor
+assert max(m.rms_to_db(v) for v in rq) < floor
+print("floor gating: PASS")
+
+# --- Stop-block handback JSON shape ---
+reason = (
+ 'The user interrupted your spoken response and said this out loud '
+ '(transcribed): "what files changed" — respond to it now.'
+)
+out = json.dumps({"decision": "block", "reason": reason})
+parsed = json.loads(out)
+assert parsed["decision"] == "block" and "transcribed" in out
+print("handback JSON: PASS")
+
+# --- recorder_command picks nothing in offline sandboxes (returns []) ---
+cmds = m.recorder_command(tmp / "reply.wav", 5)
+print("recorder candidates found offline:", cmds)
+print("ALL SPEECH EXTENSION TESTS: PASS")
diff --git a/plugins/speech/skills/voice/SKILL.md b/plugins/speech/skills/voice/SKILL.md
index 45127f6a10..d7c726eb36 100644
--- a/plugins/speech/skills/voice/SKILL.md
+++ b/plugins/speech/skills/voice/SKILL.md
@@ -68,3 +68,14 @@ Clean up the MP3 after playback. Never print or echo the API key.
The Stop-hook narrator honours three silences, checked in order:
`SPEECH_DISABLE=1`, a missing `ELEVENLABS_API_KEY`, and the flag file
`~/.sirgent/speech-disabled` (toggled by the `/hush` command).
+
+## Two-way voice (barge-in)
+
+With `/voice on` (or `SPEECH_CONVERSATION=1`) the Stop hook plays the
+response in the background and monitors the microphone. A sustained signal
+above `SPEECH_BARGE_IN_DB` (default **150**, SPL-equivalent — deliberately
+extreme so background noise never triggers it) kills playback, records up
+to `SPEECH_REPLY_SECS` seconds of the user's reply, transcribes it with the
+ElevenLabs speech-to-text API (`scribe_v1`), and hands the text back to
+SirGent by blocking the stop with the reply as the reason. The user's
+spoken words then become the next instruction.