Talk to your computer — and have it talk back — entirely on your own machine. No cloud, no account, no data leaving the room.
By Michael Amador · [email protected] · github.com/almodover
Open a web page on your phone or laptop, press Activate, and speak. HAL listens while you talk, understands when you're done, answers with a local language model and speaks the answer in a cloned voice, a fraction of a second later. You can also type.
How it works:
microphone ─▶ Voxtral Realtime (live speech-to-text)
─▶ Silero VAD + Smart Turn v3 (knows when you've finished)
─▶ your local LLM (llama.cpp or any OpenAI-compatible server)
─▶ VoxCPM2 or Pocket TTS (voice cloning, streamed)
─▶ speaker
Everything is glued together with Pipecat over WebRTC; speech models run on the GPU in audio.cpp.
- Personas: HAL is one directory in
personas/. Add your own assistant with./new-persona.sh <id> "Name"— its own prompt, voice and page, same process and port. - Optional background agent: with pi installed, HAL can hand off tasks that need the web, files or Home Assistant.
No model weights and no voice recordings are included in this repository.
| role | model | creator | license |
|---|---|---|---|
| speech-to-text | Voxtral Mini 4B Realtime 2602 | Mistral AI | Apache-2.0 |
| voice (quality) | VoxCPM2 | OpenBMB | Apache-2.0 |
| voice (fast) | Pocket TTS | Kyutai | CC-BY-4.0 |
| voice activity | Silero VAD | Silero Team | MIT |
| end of turn | Smart Turn v3 | Pipecat / Daily | BSD-2-Clause |
| language model | Ornith-1.5-35B-A3B (or any) | Ornith AI | Apache-2.0 |
| language model | Qwen3.8-Flash-Next (or any) | Qwen / Alibaba Cloud | Apache-2.0 |
The speech models are used in the GGUF format published by the audio.cpp project: audio-cpp/audio.cpp-gguf. Silero VAD and Smart Turn ship inside Pipecat.
Tested on an HP Z2 Mini G1a (AMD Ryzen AI Max+ PRO 395, Radeon 8060S,
gfx1151, 128 GB), Ubuntu 24.04, kernel 7.0. Other Ryzen AI Max / Max+
machines (Radeon 8060S / 8050S) should work the same way.
1. Give your user access to the GPU (then log out and back in):
sudo usermod -aG render,video $USER
2. Install ROCm 10 (needed to compile audio.cpp for the GPU). Add AMD's repository as described in the ROCm documentation, i.e. an apt source like:
Types: deb
URIs: https://stable.repo.amd.com/rocm/core/packages/ubuntu2404/
Suites: stable
Components: main
Signed-By: /etc/apt/keyrings/amdrocm.gpg
then install the gfx1151 development package and the build tools:
sudo apt install amdrocm-core-dev10.0-gfx1151 build-essential cmake git
3. Build audio.cpp with the HIP backend (follow its README if anything differs):
git clone https://github.com/0xShug0/audio.cpp.git ~/src/audio.cpp
cd ~/src/audio.cpp
ROCM_PATH=/opt/rocm/core-10.0 scripts/build_linux.sh --backend hip --target audiocpp_server
The server ends up in build/linux-hip-release/bin/audiocpp_server.
4. Get this repository and the speech models (~6.2 GB):
curl -LsSf https://astral.sh/uv/install.sh | sh # uv, if you don't have it
git clone https://github.com/almodover/hal9000.git ~/hal9000
cd ~/hal9000
uv sync
uvx --from huggingface_hub hf download audio-cpp/audio.cpp-gguf --local-dir models/audio \
--include "Voxtral-Mini-4B-Realtime-2602-GGUF/*q4_k.gguf" \
--include "VoxCPM2-GGUF/voxcpm2-q8_0.gguf" \
--include "PocketTTS-GGUF/english/*q8_0.gguf" \
--include "PocketTTS-GGUF/english/embeddings/*"
Point the audio.cpp config at them:
sed "s#MODELS#$PWD/models/audio#g" server.json.example > server.json
5. Start the speech server (leave it running):
LD_LIBRARY_PATH=/opt/rocm/core-10.0/lib:/opt/rocm/core-10.0/lib/llvm/lib \
~/src/audio.cpp/build/linux-hip-release/bin/audiocpp_server --config server.json --no-ui
6. Start a language model on http://127.0.0.1:8081/v1. Any
OpenAI-compatible server works; with llama.cpp
and e.g. Ornith-1.5-35B-A3B GGUF:
llama-server -m <model>.gguf --port 8081 --jinja
7. Give HAL a voice. No voice is included: record about 10 seconds of a
voice you have the rights to use (your own is simplest) as mono WAV,
16 kHz, 16-bit, save it as personas/hal/voice.wav, and write its exact
transcript in personas/hal/persona.env (HAL_VOICE_TEXT=).
8. HTTPS for the microphone. Browsers only allow the mic on localhost
or HTTPS. To talk from your phone, create a self-signed certificate once:
mkdir -p tls && openssl req -x509 -newkey rsa:2048 -nodes -days 3650 \
-keyout tls/key.pem -out tls/cert.pem -subj "/CN=hal9000"
9. Run HAL:
uv run bot.py -t webrtc --host 0.0.0.0 --port 8093
Open https://<your-machine>:8093/hal/, accept the certificate warning, press
Activate and allow the microphone. "Good afternoon, Dave."
Environment variables (per-persona values go in personas/<id>/persona.env):
| variable | default | meaning |
|---|---|---|
HAL_LLM_URL |
http://127.0.0.1:8081/v1 |
OpenAI-compatible chat endpoint |
HAL_LLM_MODEL |
local |
model name sent to it |
HAL_AUDIO_URL |
http://127.0.0.1:8083 |
audio.cpp server |
HAL_TTS_MODEL |
hal-tts-hq |
hal-tts-hq = VoxCPM2 (best voice), hal-tts = Pocket TTS (fastest) |
HAL_USER_NAME |
Dave |
what the assistant calls you |
HAL_PI |
1 |
set 0 to disable the background agent |
HAL_PI_PROVIDER / HAL_PI_MODEL |
— | provider / model from your pi configuration |
HAL_HA_URL |
— | Home Assistant URL for the background agent |
To switch a persona off without restarting: touch personas/<id>/disabled.
If pi is installed (on your PATH, or
HAL_PI_BIN), HAL gets a delegate_to_pi tool for things it can't answer from
memory: web lookups, news, files, calculations, Home Assistant. pi has shell
access to your machine — only enable it on a machine you trust it with, and
never expose port 8093 to the internet.
- HAL 9000 voice assistant — Michael Amador [email protected]
- Models: Mistral AI, OpenBMB, Kyutai, Silero Team, Pipecat / Daily, Ornith AI, Qwen — see the table above and NOTICE.
- Software: audio.cpp (ShugoAI), Pipecat (Daily), pi (Mario Zechner), llama.cpp.
HAL 9000 is a character from 2001: A Space Odyssey (1968). This is an unofficial fan project, not affiliated with or endorsed by the rights holders; it contains no audio or images from the film. HP and Z2 are trademarks of HP Inc.; AMD, Ryzen, Radeon and ROCm are trademarks of Advanced Micro Devices, Inc.
Apache-2.0. See NOTICE for credits.
Questions, bugs, ideas or just to say hi: Michael Amador — [email protected] or open an issue on GitHub.