Skip to content

Add /dictate: voice input via ElevenLabs STT (mic → prompt) - #4

Merged
nastylex merged 1 commit into
mainfrom
speech/dictate
Sep 28, 2026
Merged

nastylex merged 1 commit into
mainfrom
speech/dictate

Conversation

@freebuff-web

@freebuff-web freebuff-web Bot commented Sep 28, 2026

Copy link
Copy Markdown

Summary

Adds the input side of the speech plugin: /dictate — speak your prompt instead of typing it.

  • hooks/dictate_stt.py — records up to SPEECH_DICTATE_SECS (default 15 s) from the default microphone via sox, arecord (ALSA), or ffmpeg (macOS avfoundation), then transcribes with ElevenLabs speech-to-text (scribe_v1, override via SPEECH_STT_MODEL) and prints the recognized words to stdout as plain text.
  • commands/dictate.md — the /dictate [context] slash command: checks prerequisites (API key + a recorder) before recording, plays a listening cue, runs the bridge exactly once, and delivers the transcribed words as the user's request. Optional $ARGUMENTS frames the dictation (e.g. argument "refactor plan" + dictation "split the parser module").

Contract

  • stdout = recognized words (plain text, nothing else)
  • stderr speech: … lines = the exact reason a step failed
  • exit 1 = no key / no recorder / empty recording / API error — the command surfaces the reason and never fabricates dictation content

Fail-soft verification

  • No API key → speech: set ELEVENLABS_API_KEY… on stderr, exit 1 ✅
  • No recorder → speech: no mic recorder found. Install sox…, exit 1 ✅
  • JSON manifests still valid; Python compiles clean ✅

Pair with the existing narration (/speak, Stop-hook narrator) for full voice I/O. Base is main where PR #2's speech plugin already landed.

🤖 Generated with Codebuff

New dictate_stt.py bridge records up to SPEECH_DICTATE_SECS (15 s) from
the default mic via sox/arecord/ffmpeg, transcribes with ElevenLabs
speech-to-text (scribe_v1), and prints the words to stdout as plain text
for the /dictate command to deliver as the user's prompt.

Fails soft everywhere: missing key, missing recorder, empty recording,
or API errors print a "speech:" reason on stderr and exit 1 — the command
surfaces the reason and never fabricates content. README, plugin manifest,
and hooks.json description updated to cover the input side.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <[email protected]>
@nastylex
nastylex merged commit 779eee2 into main Sep 28, 2026
@nastylex

Copy link
Copy Markdown
Owner

easy

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants