Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,7 @@ For more installation options, uninstall steps, and troubleshooting, see the [se

## Plugins

This repository includes several plugins that extend functionality with custom commands and agents — including **speech**: voice narration of responses via ElevenLabs text-to-speech, with `/speak` and `/hush` commands. See the [plugins directory](./plugins/README.md) for detailed documentation on available plugins.
This repository includes several plugins that extend functionality with custom commands and agents — including **speech**: voice narration of responses via ElevenLabs text-to-speech, with `/speak` and `/hush` commands and two-way conversation mode (`/voice`) that lets you talk over a response to interrupt it, gated at a 150 dB-equivalent floor so background noise never triggers it. See the [plugins directory](./plugins/README.md) for detailed documentation on available plugins.

## Reporting Bugs

Expand Down
2 changes: 1 addition & 1 deletion plugins/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ Learn more in the [official plugins documentation](https://docs.claude.com/en/do
| [pr-review-toolkit](./pr-review-toolkit/) | Comprehensive PR review agents specializing in comments, tests, error handling, type design, code quality, and code simplification | **Command:** `/pr-review-toolkit:review-pr` - Run with optional review aspects (comments, tests, errors, types, code, simplify, all)<br>**Agents:** `comment-analyzer`, `pr-test-analyzer`, `silent-failure-hunter`, `type-design-analyzer`, `code-reviewer`, `code-simplifier` |
| [ralph-wiggum](./ralph-wiggum/) | Interactive self-referential AI loops for iterative development. Claude works on the same task repeatedly until completion | **Commands:** `/ralph-loop`, `/cancel-ralph` - Start/stop autonomous iteration loops<br>**Hook:** Stop - Intercepts exit attempts to continue iteration |
| [security-guidance](./security-guidance/) | Security reminder hook that warns about potential security issues when editing files | **Hook:** PreToolUse - Monitors 9 security patterns including command injection, XSS, eval usage, dangerous HTML, pickle deserialization, and os.system calls |
| [speech](./speech/) | Voice narration of responses via ElevenLabs text-to-speech | **Commands:** `/speak` - say arbitrary text aloud, `/hush` - mute/unmute the narrator<br>**Hook:** Stop - narrates SirGent's final response through the platform audio player (afplay/mpv/ffplay/PowerShell) |
| [speech](./speech/) | Voice narration and two-way conversation via ElevenLabs TTS/STT | **Commands:** `/speak` - say arbitrary text aloud, `/hush` - mute/unmute the narrator, `/voice` - toggle two-way mode<br>**Hook:** Stop - narrates the final response; with `/voice on`, talking over it (above a 150 dB-equivalent floor) interrupts playback and your transcribed reply is answered |

## Installation

Expand Down
20 changes: 6 additions & 14 deletions plugins/speech/README.md
Original file line number Diff line number Diff line change
@@ -1,19 +1,17 @@
# speech

Speech capabilities for SirGent AI: voice narration of responses through
ElevenLabs text-to-speech, voice dictation into your prompts through
ElevenLabs speech-to-text, an on-demand `/speak` command, a `/hush` mute
toggle, and a `/dictate` command for voice input.


## What it does

| Piece | Where | What happens |
| --- | --- | --- |
| Stop hook narrator | `hooks/speak_response.py` | When SirGent finishes a turn, the final assistant message is distilled (markdown stripped, capped at ~400 chars), converted to MP3 with ElevenLabs TTS, and played through the platform's audio player. |
| Two-way voice (barge-in) | same hook, `/voice on` | While the response plays, the mic is monitored. A signal above the barge-in floor lowers (kills) playback and records your reply; ElevenLabs speech-to-text transcribes it and SirGent responds to your words. |
| `/speak` command | `commands/speak.md` | Speak any text the user supplies, on demand. |
| `/hush` command | `commands/hush.md` | Mute/unmute the Stop-hook narrator with a flag file (`~/.sirgent/speech-disabled`). |
| `/dictate` command | `commands/dictate.md` + `hooks/dictate_stt.py` | Voice input: record from the default mic, transcribe with ElevenLabs STT, and deliver the words as your prompt to SirGent. |
| Voice skill | `skills/voice/SKILL.md` | Guidance for preparing natural spoken text, calling the TTS API, and playing audio on macOS/Linux/Windows. |


## Setup

Expand All @@ -30,9 +28,7 @@ toggle, and a `/dictate` command for voice input.
3. Make sure a player exists for your platform: `afplay` ships with macOS;
on Linux install `mpv` or `ffmpeg`; on Windows the PowerShell
`Media.SoundPlayer` path needs no extra install.
4. For `/dictate`, install a mic recorder: `sox` (recommended, all
platforms via Homebrew/apt/choco), or use `arecord` (ALSA Linux) /
`ffmpeg` (macOS avfoundation).


Without an API key the narrator stays silent (one setup hint on first run)
and everything else keeps working.
Expand All @@ -41,6 +37,7 @@ and everything else keeps working.

- `SPEECH_DISABLE=1` — disables the Stop-hook narrator entirely
- `/hush off` — mutes via the flag file; `/hush on` unmutes
- `/voice off` — disables mic monitoring / barge-in
- Removing the plugin disables everything

The hook always exits 0 and never blocks the session: a missing key, an API
Expand Down Expand Up @@ -72,12 +69,7 @@ speech/
├── commands/
│ ├── speak.md # /speak — say arbitrary text aloud
│ ├── hush.md # /hush — mute/unmute the narrator
│ └── dictate.md # /dictate — voice input (mic → STT → prompt)
├── hooks/
│ ├── hooks.json # Stop hook wiring
│ ├── speak_response.py # the narrator (TTS + playback)
│ ├── dictate_stt.py # dictation bridge (record → STT → stdout)
│ └── tts-python.sh # python3 finder shim (Windows-safe)

├── skills/
│ └── voice/SKILL.md # voice narration guidance
└── README.md
Expand Down
41 changes: 41 additions & 0 deletions plugins/speech/commands/voice.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
---
description: Toggle two-way voice conversation (talk over responses; SirGent hears you)
argument-hint: "[on|off]"
allowed-tools: Bash(test:*), Bash(mkdir:*), Bash(rm:*)
---

# /voice — two-way conversation mode

**Argument:** $ARGUMENTS

Two-way mode lets you interrupt SirGent's spoken responses by talking over
them (barge-in) and have your words transcribed and answered. The
interruption floor is deliberately extreme — **150 dB-equivalent by default**
— so ordinary background noise can never trigger it; only a sustained, very
loud signal (a genuine shout toward the mic) does.

## Behaviour

1. If the argument (case-insensitive) is `on`, create the flag file:
`mkdir -p ~/.sirgent && touch ~/.sirgent/speech-conversation`
Then confirm: "🎙️ Two-way voice ON. Talk over a response — loudly — to
interrupt; your words are transcribed and answered. Requires a working
`sox` or `arecord`/`ffmpeg` mic recorder. Floor: 150 dB-equivalent
(raise with SPEECH_BARGE_IN_DB)."
2. If the argument (case-insensitive) is `off`, delete it:
`rm -f ~/.sirgent/speech-conversation`
Then confirm: "🔇 Two-way voice OFF. Responses are narrated; the mic is
not monitored."
3. With no argument: check state with
`test -f ~/.sirgent/speech-conversation && echo on || echo off`
and report it plus current usage.
4. Any other value: show usage `/voice [on|off]` and stop.

## Notes

- The Stop hook honours `SPEECH_CONVERSATION=1` as well; the flag file wins
so `/voice off` always silences monitoring immediately.
- Recording needs a local recorder: `sox`, `arecord` (ALSA), or `ffmpeg`
(macOS avfoundation). Transcription reuses `ELEVENLABS_API_KEY`
(scribe model, override with `SPEECH_STT_MODEL`).
- Max reply length is `SPEECH_REPLY_SECS` (default 12 s).
16 changes: 0 additions & 16 deletions plugins/speech/hooks/hooks.json
Original file line number Diff line number Diff line change
@@ -1,17 +1 @@
{
"description": "Speech plugin — narrates SirGent's final response aloud via ElevenLabs TTS; /dictate covers the input side (mic → STT → prompt)",
"hooks": {
"Stop": [
{
"hooks": [
{
"type": "command",
"command": "bash \"${SIRGENT_PLUGIN_ROOT}/hooks/tts-python.sh\" \"${SIRGENT_PLUGIN_ROOT}/hooks/speak_response.py\"",
"asyncRewake": false,
"timeout": 60
}
]
}
]
}
}
Loading