A personal AI OS with chat, voice, memory, tool use, and remote PC control.
| Component | What it is | Where |
|---|---|---|
| ARIA server | Express server, AI routing, memory, agentic tools | server.js |
| Tools | Calc, weather, notes, todo, timer, search, news, calendar… | tools/index.js |
| Web UI | Front-end (chat, voice, settings, claw panel) | public/ |
| Claw Relay (PC) | Runs on your computer; lets ARIA control keyboard/mouse | claw-relay.js |
| ESP32 Relay | Same as above but over BLE HID for Chromebooks/sandboxed devices | ARIA_ESP32__Relay/ |
| Screenshot Watcher | Companion script for ESP32 to enable vision on a Chromebook | aria-screenshot-watcher.js |
| Voice Hook (PC) | Gives ARIA a phone number via Google Voice, password-locked | aria-voice-hook.js |
| Ollama Hook (PC) | Lets a hosted ARIA use the Ollama models on your PC | aria-ollama-hook.js |
git clone https://github.com/DaEpickid540/ARIA.git
cd ARIA
npm install
cp .env.example .env # fill in at least one provider key
npm startThen open http://localhost:3000.
At minimum you need one AI provider. See .env.example for the full list.
| Key | Required? | Purpose |
|---|---|---|
OPENROUTER_API_KEY |
recommended | Default chat model |
GROQ_API_KEY |
optional | Fast inference |
OPENAI_KEY |
optional | DALL-E image generation |
CLOUDFLARE_AI_API + CLOUDFLARE_ACCOUNT_ID |
optional | FLUX image gen |
NEWSDATA_KEY |
optional | Live news headlines |
Set ARIA_ACCESS_KEY (and ideally ARIA_RELAY_KEY) in the Render environment.
The lock screen sends what you type to POST /api/auth/login; a match sets an
HttpOnly session cookie that every /api route checks. Relays authenticate with
the relay key instead. With no key on a public deploy, the API stays open but
Claw refuses to run. Locally (no RENDER env), everything is open as before.
Claw lets ARIA control keyboard, mouse, screenshots, app launching on a target machine. Two relay flavors:
node claw-relay.js https://your-aria-url.onrender.com --key=<ARIA_RELAY_KEY>The model now gets each command's real result (shell output, errors) back
instead of "queued". Destructive shell commands (rm -rf, Remove-Item,
format, shutdown, curl … | sh, …) are held for approval no matter how the
model phrases them, as is any PC action after ARIA has read web content in the
same turn. Approvals are held server-side by id and expire after 10 minutes.
- Open
ARIA_ESP32__Relay/ARIA_ESP32_Relay.inoin Arduino IDE - Set partition scheme to Huge APP (3MB No OTA/1MB SPIFFS)
- Install libraries:
NimBLE-Arduino,ArduinoJson - Edit
WIFI_NETWORKS,SERVER_URLandRELAY_KEYconstants - Flash, then pair "ARIA Claw" from Bluetooth settings on target device
For screenshots on Chromebook (since ESP32 has no screen capture), also run:
node aria-screenshot-watcher.js https://your-aria-url.onrender.com --key=<ARIA_RELAY_KEY>This watches ~/Downloads and uploads screenshots to ARIA when the ESP32 triggers a capture.
On Render, localhost:11434 is Render's machine, so ARIA can't see the Ollama
on your PC by itself. aria-ollama-hook.js runs next to Ollama, connects out
to ARIA (nothing to port-forward) and runs model requests locally:
node aria-ollama-hook.js https://your-aria-url.onrender.com --key=<ARIA_RELAY_KEY>Your models then appear in the model switcher in the chat header; pick one there. That choice is also saved on the server, so texts through the voice hook use it too. If the hook goes offline, the switcher turns amber and ARIA falls back to a cloud model.
--ctx=16384(default) sets the context window. ARIA's system prompt is about 3k tokens, and Ollama's default on GPUs under 24 GB is 4k, which leaves no room for the conversation. Use less if the model spills out of VRAM (ollama psshould say 100% GPU);--ctx=0keeps Ollama's own setting.--ollama=http://host:11434if Ollama isn't on this machine's localhost.- Newly pulled models show up within 30 seconds, no restart needed.
aria-voice-hook.js runs on your PC, keeps voice.google.com open in a real
browser (Playwright driving your installed Edge/Chrome), and answers texts to
your Google Voice number with ARIA.
npm install # adds playwright-core; downloads no browser
# .env: ARIA_SMS_PASSWORD=<8+ chars>, plus ARIA_ACCESS_KEY if the server has one
node aria-voice-hook.js https://your-aria-url.onrender.comThe first run opens a browser window: sign in to Google Voice there. The login
is kept in data/gvoice-profile/ (your Google session, so keep it private);
after that you can add --headless.
The lock. Every conversation starts locked. Text the password to unlock it;
until then nothing gets a reply and nothing is sent on to the ARIA server.
After 5 minutes with no texts either way it locks again (--idle=<min>), and
the next text gets one "locked" notice. Text lock to lock right away. Five
texts to a locked conversation that aren't the password mute it for 15
minutes, password included. Restarting the hook locks everything.
--allow=+15551234567 limits unlocking to your own number(s).
Claw over text. If ARIA wants to run something that needs approval, it
texts you what it wants to do; reply YES or NO.
Caveats. Google Voice has no API. This drives the web page, so a Google UI
change can break it; the page selectors are all in SEL near the top of the
file. Google's Voice Acceptable Use Policy prohibits sending messages via an
automated process and can suspend numbers that break it. The hook keeps volume
low (it only answers unlocked conversations, and mutes a thread that gets 8+
replies in a minute), but that is a risk to your number. Use a Google account you can live
without.
| Endpoint | Use |
|---|---|
GET /api/health |
Server status, relay count, uptime |
POST /api/chat |
Main chat endpoint (streaming + non-streaming) |
GET /api/memory |
Read ARIA's fact memory |
POST /api/memory |
Add/delete/clear facts |
POST /api/imagine |
Image generation |
POST /api/claw/relay/register |
Relay handshake |
GET /api/claw/queue |
Relay polls commands here |
POST /api/claw/relay/result |
Relay reports command result + screenshots |
POST /api/claw/kill |
Emergency stop — clears all queues |
POST /api/auth/login |
Exchange ARIA_ACCESS_KEY for a session cookie |
POST /api/confirm |
Approve/deny a held Claw action by id |
GET/POST /api/model |
The model switcher's pick; used by requests that name no provider |
/api/ollama/relay/* |
Ollama hook: register, long-poll jobs, return results |
ARIA stores state in data/:
memory.json— facts and session historychats.json— user conversation historynotes.json,todos.json— user notes and tasksbehavior.json— adaptive personality data
All writes are atomic (temp file + rename) and debounced to avoid hammering disk.
Agentic pipeline. When ARIA needs a tool, the model emits ACTION: toolname | input. The pipeline parses this, runs the tool, injects the result back as a user message, and re-prompts up to 8 iterations.
Tool calls in the chat. Every tool call behind a reply appears in a panel above it: tool, input, result, time, status, and for web lookups the pages read. Each row expands. The panel is saved with the chat, so it survives a reload.
Web lookups read 6 pages. research and search (now the same tool) and
the fact-check agent each read 6 pages. A page that fails to load is replaced
by the next result. If the web runs out of readable results, the answer says
how many were read. scrape still reads the one URL it's given.
Sub-agents (spawn). ARIA can brief its own agents and run up to 4 in
parallel:
ACTION: spawn | name | the agent's instructions | its task
ACTION: spawn | [{"name":"for","prompt":"…","task":"…"},{"name":"against","prompt":"…","task":"…"}]
- Tools: each agent runs the same tool loop with your instructions as its system prompt, but only read-only tools: research, scrape, calc, convert, time, weather, news and the fixed agents.
- Limits: no PC control, approvals, tasks or further agents. Each agent gets 5 tool rounds and 3 minutes.
- Results: reports go back to ARIA, which writes the reply. Their tool calls appear nested under the spawn row.
Streaming. SSE-based. Chat replies stream token-by-token. If the streamed reply contains an ACTION:, the pipeline runs after the stream finishes.
Memory. Two layers: fact extraction (regex-based, runs on every reply) and behavior signals (positive/negative feedback). Both persist to data/ and inject into the system prompt on subsequent turns.
Current: Mark 1.0 (public/js/version.js is the single source of truth — bump mark and point there).
Personal project — no formal license.
Cowork/Copilot-style task engine. Tasks survive server restarts, run as multi-step plans, support scheduling, and broadcast live progress to all connected clients via SSE.
# Fire-and-forget: ARIA plans + executes
curl -X POST /api/tasks/create -H "Content-Type: application/json" -d '{
"description": "Research the top 5 open-source LLMs and summarize tradeoffs"
}'
# Plan only — review and approve before execution
curl -X POST /api/tasks/create -d '{
"description": "Refactor the Mason Navigator routing logic",
"autoExecute": false
}'
# Schedule: run once at a specific time
curl -X POST /api/tasks/create -d '{
"description": "Summarize my chats from today",
"schedule": { "runAt": 1735718400000 }
}'
# Recurring: every weekday at 9am
curl -X POST /api/tasks/create -d '{
"description": "Read overnight news and produce a 5-bullet briefing",
"schedule": { "cron": "0 9 * * 1-5" }
}'| Endpoint | Use |
|---|---|
POST /api/tasks/create |
Create a task |
GET /api/tasks |
List all tasks (optional ?status=running) |
GET /api/tasks/:id |
Get one task with full step state |
POST /api/tasks/:id/approve |
Approve a plan and start execution |
POST /api/tasks/:id/pause |
Pause a running task |
POST /api/tasks/:id/resume |
Resume a paused task |
POST /api/tasks/:id/cancel |
Cancel a task |
DELETE /api/tasks/:id |
Delete a task |
POST /api/tasks/:id/edit-steps |
Edit plan before approval |
GET /api/tasks/subscribe |
SSE feed of live updates |
GET /api/tasks/stats |
Task engine stats |
The legacy /api/background endpoints still work — they're shimmed to use the new engine.