Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions .sirgent-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,17 @@
"source": "./plugins/code-review",
"category": "productivity"
},
{
"name": "speech",
"description": "Speech capabilities for SirGent AI: narrates responses aloud via ElevenLabs text-to-speech with /speak and /hush commands",
"version": "1.0.0",
"author": {
"name": "SirGent AI",
"email": "[email protected]"
},
"source": "./plugins/speech",
"category": "productivity"
},
{
"name": "commit-commands",
"description": "Commands for git commit workflows including commit, push, and PR creation",
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,7 @@ For more installation options, uninstall steps, and troubleshooting, see the [se

## Plugins

This repository includes several plugins that extend functionality with custom commands and agents. See the [plugins directory](./plugins/README.md) for detailed documentation on available plugins.
This repository includes several plugins that extend functionality with custom commands and agents — including **speech**: voice narration of responses via ElevenLabs text-to-speech, with `/speak` and `/hush` commands. See the [plugins directory](./plugins/README.md) for detailed documentation on available plugins.

## Reporting Bugs

Expand Down
1 change: 1 addition & 0 deletions plugins/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@ Learn more in the [official plugins documentation](https://docs.claude.com/en/do
| [pr-review-toolkit](./pr-review-toolkit/) | Comprehensive PR review agents specializing in comments, tests, error handling, type design, code quality, and code simplification | **Command:** `/pr-review-toolkit:review-pr` - Run with optional review aspects (comments, tests, errors, types, code, simplify, all)<br>**Agents:** `comment-analyzer`, `pr-test-analyzer`, `silent-failure-hunter`, `type-design-analyzer`, `code-reviewer`, `code-simplifier` |
| [ralph-wiggum](./ralph-wiggum/) | Interactive self-referential AI loops for iterative development. Claude works on the same task repeatedly until completion | **Commands:** `/ralph-loop`, `/cancel-ralph` - Start/stop autonomous iteration loops<br>**Hook:** Stop - Intercepts exit attempts to continue iteration |
| [security-guidance](./security-guidance/) | Security reminder hook that warns about potential security issues when editing files | **Hook:** PreToolUse - Monitors 9 security patterns including command injection, XSS, eval usage, dangerous HTML, pickle deserialization, and os.system calls |
| [speech](./speech/) | Voice narration of responses via ElevenLabs text-to-speech | **Commands:** `/speak` - say arbitrary text aloud, `/hush` - mute/unmute the narrator<br>**Hook:** Stop - narrates SirGent's final response through the platform audio player (afplay/mpv/ffplay/PowerShell) |

## Installation

Expand Down
10 changes: 10 additions & 0 deletions plugins/speech/.sirgent-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
{
"name": "speech",
"version": "1.0.0",
"description": "Speech capabilities for SirGent AI. Narrates responses aloud through a Stop hook using ElevenLabs text-to-speech, with /speak and /hush slash commands, a voice skill, and playback via native players or afplay/mpv/powershell.",
"author": {
"name": "SirGent AI",
"email": "[email protected]"
},
"homepage": "https://github.com/sirgent-ai/sirgent-ai/tree/main/plugins/speech"
}
74 changes: 74 additions & 0 deletions plugins/speech/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
# speech

Speech capabilities for SirGent AI: voice narration of responses through
ElevenLabs text-to-speech, an on-demand `/speak` command, and a `/hush`
mute toggle.

## What it does

| Piece | Where | What happens |
| --- | --- | --- |
| Stop hook narrator | `hooks/speak_response.py` | When SirGent finishes a turn, the final assistant message is distilled (markdown stripped, capped at ~400 chars), converted to MP3 with ElevenLabs TTS, and played through the platform's audio player. |
| `/speak` command | `commands/speak.md` | Speak any text the user supplies, on demand. |
| `/hush` command | `commands/hush.md` | Mute/unmute the Stop-hook narrator with a flag file (`~/.sirgent/speech-disabled`). |
| Voice skill | `skills/voice/SKILL.md` | Guidance for preparing natural spoken text, calling the TTS API, and playing audio on macOS/Linux/Windows. |

## Setup

1. Get an API key at [elevenlabs.io](https://elevenlabs.io) (free tier
available) and set it as `ELEVENLABS_API_KEY` — in Settings →
Environment, or export it in your shell.
2. Optional voice settings:
- `ELEVENLABS_VOICE_ID` — any voice from your ElevenLabs library
(default `21m00Tcm4TlvDq8ikWAM`, "Rachel")
- `ELEVENLABS_MODEL_ID` — default `eleven_multilingual_v2`
- `SPEECH_MAX_CHARS` — narration cap, default `400`
3. Make sure a player exists for your platform: `afplay` ships with macOS;
on Linux install `mpv` or `ffmpeg`; on Windows the PowerShell
`Media.SoundPlayer` path needs no extra install.

Without an API key the narrator stays silent (one setup hint on first run)
and everything else keeps working.

## Kill switches

- `SPEECH_DISABLE=1` — disables the Stop-hook narrator entirely
- `/hush off` — mutes via the flag file; `/hush on` unmutes
- Removing the plugin disables everything

The hook always exits 0 and never blocks the session: a missing key, an API
error, or a missing player degrades to a `systemMessage` note (or silence),
never to a failed turn.

## Testing

Try it without touching hooks:

```sh
echo '{"transcript_path": ""}' | \
ELEVENLABS_API_KEY=sk-... python3 hooks/speak_response.py <<< '{}'
```

or simply run `sirgent` with the plugin loaded and let it finish a turn:

```sh
sirgent --plugin-dir plugins/speech
```

## File layout

```
speech/
├── .sirgent-plugin/
│ └── plugin.json # plugin metadata
├── commands/
│ ├── speak.md # /speak — say arbitrary text aloud
│ └── hush.md # /hush — mute/unmute the narrator
├── hooks/
│ ├── hooks.json # Stop hook wiring
│ ├── speak_response.py # the narrator (TTS + playback)
│ └── tts-python.sh # python3 finder shim (Windows-safe)
├── skills/
│ └── voice/SKILL.md # voice narration guidance
└── README.md
```
32 changes: 32 additions & 0 deletions plugins/speech/commands/hush.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
---
description: Toggle voice narration of SirGent's responses on or off
argument-hint: "[on|off]"
allowed-tools: Bash(test:*), Bash(mkdir:*), Bash(rm:*)
---

# /hush — mute or unmute voice narration

**Argument:** $ARGUMENTS

The speech plugin narrates SirGent's final responses through a Stop hook. It
is controlled by the `SPEECH_DISABLE` environment variable and a local flag
file this command manages.

## Behaviour

1. If the argument (case-insensitive) is `off`, create the flag file:
`mkdir -p ~/.sirgent && touch ~/.sirgent/speech-disabled`
Then confirm: "🔇 Voice narration muted. Run `/hush on` to unmute."
2. If the argument (case-insensitive) is `on`, delete it:
`rm -f ~/.sirgent/speech-disabled`
Then confirm: "🔊 Voice narration enabled. SirGent will speak responses
aloud (requires ELEVENLABS_API_KEY)."
3. With no argument: check whether `~/.sirgent/speech-disabled` exists
(`test -f ~/.sirgent/speech-disabled && echo muted || echo unmuted`) and
report the current state, plus a reminder of the on/off usage.
4. Treat any other value as invalid: show usage `/hush [on|off]` and stop.

Note for the user: muting here only affects this machine's flag file; the
hook also honours `SPEECH_DISABLE=1` and an unset `ELEVENLABS_API_KEY`
(always silent). The Stop hook script checks the same flag file path, so
`/hush off` wins even when the API key is configured.
44 changes: 44 additions & 0 deletions plugins/speech/commands/speak.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
---
description: Speak a message aloud through ElevenLabs text-to-speech
argument-hint: "[text to speak]"
allowed-tools: Bash(python3:*), Bash(python:*), Bash(afplay:*), Bash(mpv:*), Bash(ffplay:*), Bash(powershell:*)
---

# /speak — voice narration

Convert the user's text to speech with the ElevenLabs TTS API and play it.

**Text to speak:** $ARGUMENTS

If no text was provided, ask the user what they'd like spoken and stop.

## Steps

1. Check that the `ELEVENLABS_API_KEY` environment variable is set (run
`test -n "$ELEVENLABS_API_KEY" && echo set || echo missing`). If missing,
tell the user to add it in Settings → Environment (or export it in their
shell) and stop — do not attempt the call without a key.
2. Write the text to a temporary file `speech-input.txt` (one line, no
markdown) using a heredoc so quoting survives.
3. Call the TTS API and save MP3 output to `sirgent-speech.mp3`:

```bash
curl -sS --fail-with-body https://api.elevenlabs.io/v1/text-to-speech/21m00Tcm4TlvDq8ikWAM \
-H "xi-api-key: $ELEVENLABS_API_KEY" \
-H "Content-Type: application/json" \
-H "Accept: audio/mpeg" \
-d "{\"text\": \"$(cat speech-input.txt)\", \"model_id\": \"eleven_multilingual_v2\"}" \
-o sirgent-speech.mp3
```

Replace the voice ID in the URL with `$ELEVENLABS_VOICE_ID` when set.
4. Play the audio with the first available player for the platform:
- macOS: `afplay sirgent-speech.mp3`
- Linux: `mpv --no-video --really-quiet sirgent-speech.mp3`
(or `ffplay -nodisp -autoexit -loglevel quiet sirgent-speech.mp3`)
- Windows: `powershell -c "(New-Object Media.SoundPlayer 'sirgent-speech.mp3').PlaySync()"`
5. Delete `speech-input.txt` and `sirgent-speech.mp3` afterwards.
6. Report success or, if the API returned an error, show its message and
suggest checking the API key, voice ID, and quota.

Do not read or print the API key itself at any point.
17 changes: 17 additions & 0 deletions plugins/speech/hooks/hooks.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
{
"description": "Speech plugin — narrates SirGent's final response aloud via ElevenLabs TTS",
"hooks": {
"Stop": [
{
"hooks": [
{
"type": "command",
"command": "bash \"${SIRGENT_PLUGIN_ROOT}/hooks/tts-python.sh\" \"${SIRGENT_PLUGIN_ROOT}/hooks/speak_response.py\"",
"asyncRewake": false,
"timeout": 60
}
]
}
]
}
}
Loading