diff --git a/.sirgent-plugin/marketplace.json b/.sirgent-plugin/marketplace.json index dd23c6921a..457defeeb1 100644 --- a/.sirgent-plugin/marketplace.json +++ b/.sirgent-plugin/marketplace.json @@ -36,6 +36,17 @@ "source": "./plugins/code-review", "category": "productivity" }, + { + "name": "speech", + "description": "Speech capabilities for SirGent AI: narrates responses aloud via ElevenLabs text-to-speech with /speak and /hush commands", + "version": "1.0.0", + "author": { + "name": "SirGent AI", + "email": "support@sirgent.ai" + }, + "source": "./plugins/speech", + "category": "productivity" + }, { "name": "commit-commands", "description": "Commands for git commit workflows including commit, push, and PR creation", diff --git a/README.md b/README.md index 6f0949ad3c..faf80072f9 100644 --- a/README.md +++ b/README.md @@ -47,7 +47,7 @@ For more installation options, uninstall steps, and troubleshooting, see the [se ## Plugins -This repository includes several plugins that extend functionality with custom commands and agents. See the [plugins directory](./plugins/README.md) for detailed documentation on available plugins. +This repository includes several plugins that extend functionality with custom commands and agents — including **speech**: voice narration of responses via ElevenLabs text-to-speech, with `/speak` and `/hush` commands. See the [plugins directory](./plugins/README.md) for detailed documentation on available plugins. ## Reporting Bugs diff --git a/plugins/README.md b/plugins/README.md index ac89d6f06e..884411d914 100644 --- a/plugins/README.md +++ b/plugins/README.md @@ -25,6 +25,7 @@ Learn more in the [official plugins documentation](https://docs.claude.com/en/do | [pr-review-toolkit](./pr-review-toolkit/) | Comprehensive PR review agents specializing in comments, tests, error handling, type design, code quality, and code simplification | **Command:** `/pr-review-toolkit:review-pr` - Run with optional review aspects (comments, tests, errors, types, code, simplify, all)
**Agents:** `comment-analyzer`, `pr-test-analyzer`, `silent-failure-hunter`, `type-design-analyzer`, `code-reviewer`, `code-simplifier` | | [ralph-wiggum](./ralph-wiggum/) | Interactive self-referential AI loops for iterative development. Claude works on the same task repeatedly until completion | **Commands:** `/ralph-loop`, `/cancel-ralph` - Start/stop autonomous iteration loops
**Hook:** Stop - Intercepts exit attempts to continue iteration | | [security-guidance](./security-guidance/) | Security reminder hook that warns about potential security issues when editing files | **Hook:** PreToolUse - Monitors 9 security patterns including command injection, XSS, eval usage, dangerous HTML, pickle deserialization, and os.system calls | +| [speech](./speech/) | Voice narration of responses via ElevenLabs text-to-speech | **Commands:** `/speak` - say arbitrary text aloud, `/hush` - mute/unmute the narrator
**Hook:** Stop - narrates SirGent's final response through the platform audio player (afplay/mpv/ffplay/PowerShell) | ## Installation diff --git a/plugins/speech/.sirgent-plugin/plugin.json b/plugins/speech/.sirgent-plugin/plugin.json new file mode 100644 index 0000000000..61f53aec4a --- /dev/null +++ b/plugins/speech/.sirgent-plugin/plugin.json @@ -0,0 +1,10 @@ +{ + "name": "speech", + "version": "1.0.0", + "description": "Speech capabilities for SirGent AI. Narrates responses aloud through a Stop hook using ElevenLabs text-to-speech, with /speak and /hush slash commands, a voice skill, and playback via native players or afplay/mpv/powershell.", + "author": { + "name": "SirGent AI", + "email": "support@sirgent.ai" + }, + "homepage": "https://github.com/sirgent-ai/sirgent-ai/tree/main/plugins/speech" +} diff --git a/plugins/speech/README.md b/plugins/speech/README.md new file mode 100644 index 0000000000..a2686bc17a --- /dev/null +++ b/plugins/speech/README.md @@ -0,0 +1,74 @@ +# speech + +Speech capabilities for SirGent AI: voice narration of responses through +ElevenLabs text-to-speech, an on-demand `/speak` command, and a `/hush` +mute toggle. + +## What it does + +| Piece | Where | What happens | +| --- | --- | --- | +| Stop hook narrator | `hooks/speak_response.py` | When SirGent finishes a turn, the final assistant message is distilled (markdown stripped, capped at ~400 chars), converted to MP3 with ElevenLabs TTS, and played through the platform's audio player. | +| `/speak` command | `commands/speak.md` | Speak any text the user supplies, on demand. | +| `/hush` command | `commands/hush.md` | Mute/unmute the Stop-hook narrator with a flag file (`~/.sirgent/speech-disabled`). | +| Voice skill | `skills/voice/SKILL.md` | Guidance for preparing natural spoken text, calling the TTS API, and playing audio on macOS/Linux/Windows. | + +## Setup + +1. Get an API key at [elevenlabs.io](https://elevenlabs.io) (free tier + available) and set it as `ELEVENLABS_API_KEY` — in Settings → + Environment, or export it in your shell. +2. Optional voice settings: + - `ELEVENLABS_VOICE_ID` — any voice from your ElevenLabs library + (default `21m00Tcm4TlvDq8ikWAM`, "Rachel") + - `ELEVENLABS_MODEL_ID` — default `eleven_multilingual_v2` + - `SPEECH_MAX_CHARS` — narration cap, default `400` +3. Make sure a player exists for your platform: `afplay` ships with macOS; + on Linux install `mpv` or `ffmpeg`; on Windows the PowerShell + `Media.SoundPlayer` path needs no extra install. + +Without an API key the narrator stays silent (one setup hint on first run) +and everything else keeps working. + +## Kill switches + +- `SPEECH_DISABLE=1` — disables the Stop-hook narrator entirely +- `/hush off` — mutes via the flag file; `/hush on` unmutes +- Removing the plugin disables everything + +The hook always exits 0 and never blocks the session: a missing key, an API +error, or a missing player degrades to a `systemMessage` note (or silence), +never to a failed turn. + +## Testing + +Try it without touching hooks: + +```sh +echo '{"transcript_path": ""}' | \ + ELEVENLABS_API_KEY=sk-... python3 hooks/speak_response.py <<< '{}' +``` + +or simply run `sirgent` with the plugin loaded and let it finish a turn: + +```sh +sirgent --plugin-dir plugins/speech +``` + +## File layout + +``` +speech/ +├── .sirgent-plugin/ +│ └── plugin.json # plugin metadata +├── commands/ +│ ├── speak.md # /speak — say arbitrary text aloud +│ └── hush.md # /hush — mute/unmute the narrator +├── hooks/ +│ ├── hooks.json # Stop hook wiring +│ ├── speak_response.py # the narrator (TTS + playback) +│ └── tts-python.sh # python3 finder shim (Windows-safe) +├── skills/ +│ └── voice/SKILL.md # voice narration guidance +└── README.md +``` diff --git a/plugins/speech/commands/hush.md b/plugins/speech/commands/hush.md new file mode 100644 index 0000000000..5f708009b2 --- /dev/null +++ b/plugins/speech/commands/hush.md @@ -0,0 +1,32 @@ +--- +description: Toggle voice narration of SirGent's responses on or off +argument-hint: "[on|off]" +allowed-tools: Bash(test:*), Bash(mkdir:*), Bash(rm:*) +--- + +# /hush — mute or unmute voice narration + +**Argument:** $ARGUMENTS + +The speech plugin narrates SirGent's final responses through a Stop hook. It +is controlled by the `SPEECH_DISABLE` environment variable and a local flag +file this command manages. + +## Behaviour + +1. If the argument (case-insensitive) is `off`, create the flag file: + `mkdir -p ~/.sirgent && touch ~/.sirgent/speech-disabled` + Then confirm: "🔇 Voice narration muted. Run `/hush on` to unmute." +2. If the argument (case-insensitive) is `on`, delete it: + `rm -f ~/.sirgent/speech-disabled` + Then confirm: "🔊 Voice narration enabled. SirGent will speak responses + aloud (requires ELEVENLABS_API_KEY)." +3. With no argument: check whether `~/.sirgent/speech-disabled` exists + (`test -f ~/.sirgent/speech-disabled && echo muted || echo unmuted`) and + report the current state, plus a reminder of the on/off usage. +4. Treat any other value as invalid: show usage `/hush [on|off]` and stop. + +Note for the user: muting here only affects this machine's flag file; the +hook also honours `SPEECH_DISABLE=1` and an unset `ELEVENLABS_API_KEY` +(always silent). The Stop hook script checks the same flag file path, so +`/hush off` wins even when the API key is configured. diff --git a/plugins/speech/commands/speak.md b/plugins/speech/commands/speak.md new file mode 100644 index 0000000000..0fd406c16d --- /dev/null +++ b/plugins/speech/commands/speak.md @@ -0,0 +1,44 @@ +--- +description: Speak a message aloud through ElevenLabs text-to-speech +argument-hint: "[text to speak]" +allowed-tools: Bash(python3:*), Bash(python:*), Bash(afplay:*), Bash(mpv:*), Bash(ffplay:*), Bash(powershell:*) +--- + +# /speak — voice narration + +Convert the user's text to speech with the ElevenLabs TTS API and play it. + +**Text to speak:** $ARGUMENTS + +If no text was provided, ask the user what they'd like spoken and stop. + +## Steps + +1. Check that the `ELEVENLABS_API_KEY` environment variable is set (run + `test -n "$ELEVENLABS_API_KEY" && echo set || echo missing`). If missing, + tell the user to add it in Settings → Environment (or export it in their + shell) and stop — do not attempt the call without a key. +2. Write the text to a temporary file `speech-input.txt` (one line, no + markdown) using a heredoc so quoting survives. +3. Call the TTS API and save MP3 output to `sirgent-speech.mp3`: + + ```bash + curl -sS --fail-with-body https://api.elevenlabs.io/v1/text-to-speech/21m00Tcm4TlvDq8ikWAM \ + -H "xi-api-key: $ELEVENLABS_API_KEY" \ + -H "Content-Type: application/json" \ + -H "Accept: audio/mpeg" \ + -d "{\"text\": \"$(cat speech-input.txt)\", \"model_id\": \"eleven_multilingual_v2\"}" \ + -o sirgent-speech.mp3 + ``` + + Replace the voice ID in the URL with `$ELEVENLABS_VOICE_ID` when set. +4. Play the audio with the first available player for the platform: + - macOS: `afplay sirgent-speech.mp3` + - Linux: `mpv --no-video --really-quiet sirgent-speech.mp3` + (or `ffplay -nodisp -autoexit -loglevel quiet sirgent-speech.mp3`) + - Windows: `powershell -c "(New-Object Media.SoundPlayer 'sirgent-speech.mp3').PlaySync()"` +5. Delete `speech-input.txt` and `sirgent-speech.mp3` afterwards. +6. Report success or, if the API returned an error, show its message and + suggest checking the API key, voice ID, and quota. + +Do not read or print the API key itself at any point. diff --git a/plugins/speech/hooks/hooks.json b/plugins/speech/hooks/hooks.json new file mode 100644 index 0000000000..fb192aadf0 --- /dev/null +++ b/plugins/speech/hooks/hooks.json @@ -0,0 +1,17 @@ +{ + "description": "Speech plugin — narrates SirGent's final response aloud via ElevenLabs TTS", + "hooks": { + "Stop": [ + { + "hooks": [ + { + "type": "command", + "command": "bash \"${SIRGENT_PLUGIN_ROOT}/hooks/tts-python.sh\" \"${SIRGENT_PLUGIN_ROOT}/hooks/speak_response.py\"", + "asyncRewake": false, + "timeout": 60 + } + ] + } + ] + } +} diff --git a/plugins/speech/hooks/speak_response.py b/plugins/speech/hooks/speak_response.py new file mode 100644 index 0000000000..bf624fda26 --- /dev/null +++ b/plugins/speech/hooks/speak_response.py @@ -0,0 +1,227 @@ +#!/usr/bin/env python3 +""" +Speech plugin for SirGent AI — Stop hook narrator. + +Reads SirGent's final response from the Stop hook stdin payload, distills it +into a short spoken summary, converts it to speech with the ElevenLabs +text-to-speech API, and plays the audio through whatever player the platform +offers (afplay on macOS, mpv/ffplay/paplay/aplay on Linux, PowerShell +Win32 SoundPlayer on Windows). + +Configuration (environment variables): +- ELEVENLABS_API_KEY Required for cloud TTS. Get one at https://elevenlabs.io +- ELEVENLABS_VOICE_ID Voice to use. Default: 21m00Tcm4TlvDq8ikWAM ("Rachel") +- ELEVENLABS_MODEL_ID Model. Default: eleven_multilingual_v2 +- SPEECH_MAX_CHARS Truncate narration at N characters. Default: 400 +- SPEECH_DISABLE "1" fully disables the plugin +- SPEECH_AUDIO_DIR Where to cache audio files. Default: system temp + +Kill-switch behaviour: with SPEECH_DISABLE=1 or no API key, the hook exits 0 +silently so the session is never blocked. +""" + +import json +import os +import re +import subprocess +import sys +import tempfile +from pathlib import Path + +API_URL = "https://api.elevenlabs.io/v1/text-to-speech/{voice_id}" + +DEFAULT_VOICE_ID = "21m00Tcm4TlvDq8ikWAM" # "Rachel" +DEFAULT_MODEL_ID = "eleven_multilingual_v2" + +# Markdown and noise that should never be spoken aloud +MD_PATTERNS = [ + re.compile(r"```.*?```", re.DOTALL), # fenced code blocks + re.compile(r"`[^`\n]+`"), # inline code + re.compile(r"!\[[^\]]*\]\([^)]*\)"), # images + re.compile(r"\[([^\]]+)\]\([^)]*\)"), # links -> keep text + re.compile(r"^\s{0,3}#{1,6}\s+", re.MULTILINE), # headings + re.compile(r"^\s*[-*+]\s+", re.MULTILINE), # bullets + re.compile(r"\|"), # table pipes + re.compile(r"^---+\s*$", re.MULTILINE), # hrules +] + +MAX_CHARS_DEFAULT = 400 + + +def env_flag(name: str) -> bool: + return os.environ.get(name, "").strip() == "1" + + +def clean_for_speech(text: str) -> str: + """Strip markdown/noise and collapse whitespace so text reads naturally.""" + for pattern in MD_PATTERNS: + text = pattern.sub(" " if pattern.pattern != r"\[([^\]]+)\]\([^)]*\)" else r"\1", text) + text = re.sub(r"\s+", " ", text) + return text.strip() + + +def summarize(text: str, limit: int) -> str: + """Trim to the last complete sentence that fits the limit.""" + if len(text) <= limit: + return text + clipped = text[:limit] + # Prefer cutting at sentence end; fall back to word boundary. + for sep in (". ", "! ", "? ", " "): + idx = clipped.rfind(sep) + if idx > limit * 0.4: + return clipped[: idx + 1].strip() + return clipped.strip() + + +def extract_response(payload: dict) -> str: + """Pull the transcript's final assistant message out of the Stop payload.""" + transcript = payload.get("transcript_path") + messages = [] + if transcript and os.path.exists(transcript): + try: + with open(transcript, "r", encoding="utf-8", errors="replace") as fh: + for line in fh: + line = line.strip() + if not line: + continue + try: + entry = json.loads(line) + except json.JSONDecodeError: + continue + if entry.get("type") == "assistant": + msg = entry.get("message") or {} + content = msg.get("content") + if isinstance(content, str): + messages.append(content) + elif isinstance(content, list): + for block in content: + if isinstance(block, dict) and block.get("type") == "text": + messages.append(block.get("text", "")) + except OSError: + pass + if not messages: + # Fallback: hook payloads may carry last_message directly. + messages = [payload.get("last_message") or payload.get("stop_hook_active") and "" or ""] + return messages[-1] if messages else "" + + +def synthesize(text: str, api_key: str, out_path: Path) -> bool: + """POST to ElevenLabs TTS and write MP3 bytes. Returns True on success.""" + import urllib.error + import urllib.request + + voice_id = os.environ.get("ELEVENLABS_VOICE_ID", DEFAULT_VOICE_ID) + model_id = os.environ.get("ELEVENLABS_MODEL_ID", DEFAULT_MODEL_ID) + url = API_URL.format(voice_id=voice_id) + body = json.dumps({ + "text": text, + "model_id": model_id, + "voice_settings": {"stability": 0.5, "similarity_boost": 0.75}, + }).encode("utf-8") + + req = urllib.request.Request(url, data=body, method="POST", headers={ + "xi-api-key": api_key, + "Content-Type": "application/json", + "Accept": "audio/mpeg", + }) + try: + with urllib.request.urlopen(req, timeout=45) as resp: + out_path.write_bytes(resp.read()) + return out_path.stat().st_size > 0 + except urllib.error.HTTPError as exc: + detail = exc.read().decode("utf-8", "replace")[:200] + print(json.dumps({"systemMessage": f"speech: ElevenLabs HTTP {exc.code}: {detail}"})) + return False + except (urllib.error.URLError, OSError, TimeoutError) as exc: + print(json.dumps({"systemMessage": f"speech: TTS request failed: {exc}"})) + return False + + +def player_commands(mp3: Path) -> list: + """Ordered player candidates for the current platform.""" + plat = sys.platform + if plat == "darwin": + return [["afplay", str(mp3)]] + if plat.startswith("win"): + ps = ( + "(New-Object Media.SoundPlayer " + f"'{mp3}').PlaySync()" + ) + return [["powershell", "-NoProfile", "-Command", ps]] + # Linux & friends: try the common CLI players. mpv/ffplay handle MP3 best. + return [ + ["mpv", "--no-video", "--really-quiet", str(mp3)], + ["ffplay", "-nodisp", "-autoexit", "-loglevel", "quiet", str(mp3)], + ["paplay", str(mp3)], + ["aplay", str(mp3)], + ["play", str(mp3)], + ] + + +def play(mp3: Path) -> bool: + for cmd in player_commands(mp3): + try: + proc = subprocess.run( + cmd, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL, timeout=120 + ) + if proc.returncode == 0: + return True + except (OSError, subprocess.TimeoutExpired): + continue + print(json.dumps({ + "systemMessage": ( + "speech: generated audio but found no player. " + "Install mpv or ffplay, then retry." + ) + })) + return False + + +def main() -> None: + try: + payload = json.load(sys.stdin) + except json.JSONDecodeError: + payload = {} + + if env_flag("SPEECH_DISABLE"): + sys.exit(0) + + api_key = os.environ.get("ELEVENLABS_API_KEY", "").strip() + if not api_key: + # First-run hint, once, then stay quiet. + marker = Path(tempfile.gettempdir()) / "sirgent-speech-hint" + if not marker.exists(): + try: + marker.write_text("1", encoding="utf-8") + except OSError: + pass + print(json.dumps({ + "systemMessage": ( + "speech plugin: set ELEVENLABS_API_KEY to enable voice " + "narration of responses (https://elevenlabs.io). " + "Set SPEECH_DISABLE=1 to hide this hint." + ) + })) + sys.exit(0) + + raw = extract_response(payload) + if not raw.strip(): + sys.exit(0) + + spoken = summarize(clean_for_speech(raw), int(os.environ.get("SPEECH_MAX_CHARS", MAX_CHARS_DEFAULT))) + if not spoken: + sys.exit(0) + + audio_dir = Path(os.environ.get("SPEECH_AUDIO_DIR", tempfile.gettempdir())) / "sirgent-speech" + try: + audio_dir.mkdir(parents=True, exist_ok=True) + mp3 = audio_dir / "response.mp3" + if synthesize(spoken, api_key, mp3) and play(mp3): + sys.exit(0) + except OSError as exc: + print(json.dumps({"systemMessage": f"speech: {exc}"})) + sys.exit(0) + + +if __name__ == "__main__": + main() diff --git a/plugins/speech/hooks/tts-python.sh b/plugins/speech/hooks/tts-python.sh new file mode 100644 index 0000000000..6cdd92c0a9 --- /dev/null +++ b/plugins/speech/hooks/tts-python.sh @@ -0,0 +1,30 @@ +#!/usr/bin/env bash +# Find a working Python 3 interpreter and exec the hook with it. +# +# Mirrors security-guidance's sg-python.sh: on Windows + Git Bash, `python3` +# can resolve to the Microsoft Store stub that exits 49 silently. Probe each +# candidate and skip any that fails or is Python 2. +# +# Args after the shim path are passed straight through to the interpreter. +set -e + +probe() { + # $1..N: the interpreter command (may be multi-word like `py -3`) + # Probe writes the major version to stdout and exits 0 iff it's >=3. + "$@" -c 'import sys; print(sys.version_info[0])' 2>/dev/null +} + +for cmd in "python3" "python" "py -3"; do + # Word-split intentionally so `py -3` works + # shellcheck disable=SC2086 + v=$(probe $cmd) || continue + if [ "$v" = "3" ]; then + # shellcheck disable=SC2086 + exec $cmd "$@" + fi +done + +echo "speech: no working Python 3 interpreter found." >&2 +echo " tried: python3, python, py -3" >&2 +echo " on Windows, install Python from https://python.org (NOT the Microsoft Store)" >&2 +exit 0 diff --git a/plugins/speech/skills/voice/SKILL.md b/plugins/speech/skills/voice/SKILL.md new file mode 100644 index 0000000000..45127f6a10 --- /dev/null +++ b/plugins/speech/skills/voice/SKILL.md @@ -0,0 +1,70 @@ +--- +name: sirgent-voice +description: Narrate content aloud with text-to-speech. Use when the user asks to speak, say, or read aloud a message, summary, or response, or asks how voice output is configured. +--- + +# SirGent voice (text-to-speech) + +Narrate text aloud through the ElevenLabs TTS API. This skill backs the +`/speak` command and the Stop-hook narrator in the speech plugin. + +## When to use + +- The user says "say that out loud", "read it to me", "speak it" +- The user asks for a spoken summary of a response or document +- The user asks how voice narration works or how to configure it + +## Requirements + +- `ELEVENLABS_API_KEY` environment variable (get one at https://elevenlabs.io) +- Optional: `ELEVENLABS_VOICE_ID` (default voice: "Rachel", + `21m00Tcm4TlvDq8ikWAM`), `ELEVENLABS_MODEL_ID` + (default `eleven_multilingual_v2`) +- A local audio player: `afplay` (macOS), `mpv`/`ffplay` (Linux), + PowerShell `Media.SoundPlayer` (Windows) + +If the API key is missing, tell the user to add it (Settings → Environment or +shell export) rather than failing silently. + +## Preparing text to speak + +1. Strip markdown: remove code fences, inline code backticks, images, table + pipes, and heading marks; keep link text without the URL. +2. Collapse whitespace into single spaces. +3. Cap at roughly 400 characters (about 25 seconds of speech) — end at a + sentence boundary when possible. Long output should be summarized, not + truncated mid-word. + +## Calling the API + +```bash +curl -sS --fail-with-body \ + "https://api.elevenlabs.io/v1/text-to-speech/${ELEVENLABS_VOICE_ID:-21m00Tcm4TlvDq8ikWAM}" \ + -H "xi-api-key: $ELEVENLABS_API_KEY" \ + -H "Content-Type: application/json" \ + -H "Accept: audio/mpeg" \ + -d '{"text": "PREPARED_TEXT_HERE", "model_id": "eleven_multilingual_v2"}' \ + -o sirgent-speech.mp3 +``` + +Errors surface as non-zero exit or an HTTP error body: check the key first, +then the voice ID, then plan quota. + +## Playing audio + +Pick the first available player for the platform: + +| Platform | Command | +|---|---| +| macOS | `afplay sirgent-speech.mp3` | +| Linux | `mpv --no-video --really-quiet sirgent-speech.mp3` | +| Linux (fallback) | `ffplay -nodisp -autoexit -loglevel quiet sirgent-speech.mp3` | +| Windows | `powershell -c "(New-Object Media.SoundPlayer 'sirgent-speech.mp3').PlaySync()"` | + +Clean up the MP3 after playback. Never print or echo the API key. + +## Muting + +The Stop-hook narrator honours three silences, checked in order: +`SPEECH_DISABLE=1`, a missing `ELEVENLABS_API_KEY`, and the flag file +`~/.sirgent/speech-disabled` (toggled by the `/hush` command).