Skip to content

Repository files navigation

🔴 HAL 9000 — a local voice assistant

Talk to your computer — and have it talk back — entirely on your own machine. No cloud, no account, no data leaving the room.

By Michael Amador · [email protected] · github.com/almodover

Open a web page on your phone or laptop, press Activate, and speak. HAL listens while you talk, understands when you're done, answers with a local language model and speaks the answer in a cloned voice, a fraction of a second later. You can also type.

How it works:

microphone ─▶ Voxtral Realtime (live speech-to-text)
           ─▶ Silero VAD + Smart Turn v3 (knows when you've finished)
           ─▶ your local LLM (llama.cpp or any OpenAI-compatible server)
           ─▶ VoxCPM2 or Pocket TTS (voice cloning, streamed)
           ─▶ speaker

Everything is glued together with Pipecat over WebRTC; speech models run on the GPU in audio.cpp.

  • Personas: HAL is one directory in personas/. Add your own assistant with ./new-persona.sh <id> "Name" — its own prompt, voice and page, same process and port.
  • Optional background agent: with pi installed, HAL can hand off tasks that need the web, files or Home Assistant.

The models and their creators

No model weights and no voice recordings are included in this repository.

role model creator license
speech-to-text Voxtral Mini 4B Realtime 2602 Mistral AI Apache-2.0
voice (quality) VoxCPM2 OpenBMB Apache-2.0
voice (fast) Pocket TTS Kyutai CC-BY-4.0
voice activity Silero VAD Silero Team MIT
end of turn Smart Turn v3 Pipecat / Daily BSD-2-Clause
language model Ornith-1.5-35B-A3B (or any) Ornith AI Apache-2.0
language model Qwen3.8-Flash-Next (or any) Qwen / Alibaba Cloud Apache-2.0

The speech models are used in the GGUF format published by the audio.cpp project: audio-cpp/audio.cpp-gguf. Silero VAD and Smart Turn ship inside Pipecat.

Install on a Strix Halo machine

Tested on an HP Z2 Mini G1a (AMD Ryzen AI Max+ PRO 395, Radeon 8060S, gfx1151, 128 GB), Ubuntu 24.04, kernel 7.0. Other Ryzen AI Max / Max+ machines (Radeon 8060S / 8050S) should work the same way.

1. Give your user access to the GPU (then log out and back in):

sudo usermod -aG render,video $USER

2. Install ROCm 10 (needed to compile audio.cpp for the GPU). Add AMD's repository as described in the ROCm documentation, i.e. an apt source like:

Types: deb
URIs: https://stable.repo.amd.com/rocm/core/packages/ubuntu2404/
Suites: stable
Components: main
Signed-By: /etc/apt/keyrings/amdrocm.gpg

then install the gfx1151 development package and the build tools:

sudo apt install amdrocm-core-dev10.0-gfx1151 build-essential cmake git

3. Build audio.cpp with the HIP backend (follow its README if anything differs):

git clone https://github.com/0xShug0/audio.cpp.git ~/src/audio.cpp
cd ~/src/audio.cpp
ROCM_PATH=/opt/rocm/core-10.0 scripts/build_linux.sh --backend hip --target audiocpp_server

The server ends up in build/linux-hip-release/bin/audiocpp_server.

4. Get this repository and the speech models (~6.2 GB):

curl -LsSf https://astral.sh/uv/install.sh | sh        # uv, if you don't have it
git clone https://github.com/almodover/hal9000.git ~/hal9000
cd ~/hal9000
uv sync
uvx --from huggingface_hub hf download audio-cpp/audio.cpp-gguf --local-dir models/audio \
  --include "Voxtral-Mini-4B-Realtime-2602-GGUF/*q4_k.gguf" \
  --include "VoxCPM2-GGUF/voxcpm2-q8_0.gguf" \
  --include "PocketTTS-GGUF/english/*q8_0.gguf" \
  --include "PocketTTS-GGUF/english/embeddings/*"

Point the audio.cpp config at them:

sed "s#MODELS#$PWD/models/audio#g" server.json.example > server.json

5. Start the speech server (leave it running):

LD_LIBRARY_PATH=/opt/rocm/core-10.0/lib:/opt/rocm/core-10.0/lib/llvm/lib \
  ~/src/audio.cpp/build/linux-hip-release/bin/audiocpp_server --config server.json --no-ui

6. Start a language model on http://127.0.0.1:8081/v1. Any OpenAI-compatible server works; with llama.cpp and e.g. Ornith-1.5-35B-A3B GGUF:

llama-server -m <model>.gguf --port 8081 --jinja

7. Give HAL a voice. No voice is included: record about 10 seconds of a voice you have the rights to use (your own is simplest) as mono WAV, 16 kHz, 16-bit, save it as personas/hal/voice.wav, and write its exact transcript in personas/hal/persona.env (HAL_VOICE_TEXT=).

8. HTTPS for the microphone. Browsers only allow the mic on localhost or HTTPS. To talk from your phone, create a self-signed certificate once:

mkdir -p tls && openssl req -x509 -newkey rsa:2048 -nodes -days 3650 \
  -keyout tls/key.pem -out tls/cert.pem -subj "/CN=hal9000"

9. Run HAL:

uv run bot.py -t webrtc --host 0.0.0.0 --port 8093

Open https://<your-machine>:8093/hal/, accept the certificate warning, press Activate and allow the microphone. "Good afternoon, Dave."

Settings

Environment variables (per-persona values go in personas/<id>/persona.env):

variable default meaning
HAL_LLM_URL http://127.0.0.1:8081/v1 OpenAI-compatible chat endpoint
HAL_LLM_MODEL local model name sent to it
HAL_AUDIO_URL http://127.0.0.1:8083 audio.cpp server
HAL_TTS_MODEL hal-tts-hq hal-tts-hq = VoxCPM2 (best voice), hal-tts = Pocket TTS (fastest)
HAL_USER_NAME Dave what the assistant calls you
HAL_PI 1 set 0 to disable the background agent
HAL_PI_PROVIDER / HAL_PI_MODEL — provider / model from your pi configuration
HAL_HA_URL — Home Assistant URL for the background agent

To switch a persona off without restarting: touch personas/<id>/disabled.

Optional: the background agent

If pi is installed (on your PATH, or HAL_PI_BIN), HAL gets a delegate_to_pi tool for things it can't answer from memory: web lookups, news, files, calculations, Home Assistant. pi has shell access to your machine — only enable it on a machine you trust it with, and never expose port 8093 to the internet.

Credits

  • HAL 9000 voice assistant — Michael Amador [email protected]
  • Models: Mistral AI, OpenBMB, Kyutai, Silero Team, Pipecat / Daily, Ornith AI, Qwen — see the table above and NOTICE.
  • Software: audio.cpp (ShugoAI), Pipecat (Daily), pi (Mario Zechner), llama.cpp.

HAL 9000 is a character from 2001: A Space Odyssey (1968). This is an unofficial fan project, not affiliated with or endorsed by the rights holders; it contains no audio or images from the film. HP and Z2 are trademarks of HP Inc.; AMD, Ryzen, Radeon and ROCm are trademarks of Advanced Micro Devices, Inc.

License

Apache-2.0. See NOTICE for credits.

Contact

Questions, bugs, ideas or just to say hi: Michael Amador — [email protected] or open an issue on GitHub.

About

Local, private voice assistant with the HAL 9000 persona for AMD Strix Halo (ROCm 10): Voxtral Realtime STT, VoxCPM2 / Pocket TTS, Smart Turn, any local LLM, via Pipecat + audio.cpp. No models or voices included.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages