Local-first vectorized memory for AI agents. Memories are stored in repo-scoped
namespaces derived from your git origin URL, so each repository gets its own
isolated memory and switching projects switches context automatically. Recall
and embeddings run fully local — embeddings are always computed locally via
Ollama, and no memory content ever leaves your machine.
Two surfaces share one store (~/.claude/kairos/memory/ns-<namespace>.db):
- Claude Code is registered over stdio. Each session spawns
backant-memory serve, which resolves the repo-scoped store for that session's cwd — so isolation and sharing withbackant-kairosare exact per project. - A launchd-supervised daemon (label
io.backant.memory,127.0.0.1:41414) stays always-on to keep Ollama warm, answer the authenticated SessionStart/digestwarm path, and expose MCP over streamable HTTP for other agents. Note: the HTTP/mcpsurface serves a single global store (repo"") — HTTP sessions carry no cwd, so it is not repo-scoped. Roots-based HTTP scoping is a tracked follow-up; until it lands, repo isolation is delivered on the stdio path (Claude Code) only.
- Node.js >= 20
- macOS (launchd) for the always-on background service. The stdio transport
(
backant-memory serve --stdio) works on any platform for other MCP clients. - Docker (for the local Ollama embedding runtime — auto-started when needed).
npm install -g backant-memory
backant-memory installinstall is idempotent — re-running it is safe and repairs drift.
- Always-on service. Generates
~/Library/LaunchAgents/io.backant.memory.plistand bootstraps it via launchd (KeepAlive+RunAtLoad), plus a0600bearer-token file. The daemon survives sleep, crashes, and reboot. It keeps Ollama warm and serves the authenticated/digest+ HTTP/mcpsurfaces. - MCP registration (user scope, stdio). Registers the
backant-memoryserver in~/.claude.jsonas a stdio entry (command: <install>/bin/backant-memory.js,args: ["serve"]). Every Claude Code session in every repo then sees the memory tools, each session repo-scoped to its own cwd. (Other MCP clients use the HTTP endpoint — see Other MCP clients.) - Global CLAUDE.md block. Appends/updates a managed section in
~/.claude/CLAUDE.mdbetween<!-- backant-memory:start -->and<!-- backant-memory:end -->markers (content outside the markers is never touched). It nudges agents tomemory_recallbefore acting on a known topic andmemory_write_stm/memory_write_episodeafter a verified outcome. - Skill. Installs
~/.claude/skills/backant-memory/SKILL.mddescribing when each tool group applies. - Hooks in
~/.claude/settings.json(pass--no-hookto skip all three):- SessionStart — opens every session with a digest: the latest handoff brief, the latest automatic session summary ("Last session — resume here"), and a recall of durable repo knowledge.
- UserPromptSubmit — ambient recall: every prompt is used as a cue and
the top hits are injected as
## Memory recall — <repo>with tier, type, age and id. Skips slash commands and trivial prompts; never repeats a hit within a session; hard 2.5s deadline, warm path via the daemon's/recall. - PreCompact + SessionEnd — writes one deterministic
session_summaryrow per session (prompts, files touched, outcome, branch) from the transcript, in a detached worker so exit is never delayed. No model call.
The MCP entry is registered with "alwaysLoad": true so the memory tools are
never deferred behind Claude Code's tool search — a session that has to
ToolSearch for its memory tools mostly won't.
Two profiles, one implementation:
core(default for stdio / Claude Code) — nine trigger-first tools:memory_recall(cue,id, orwith_edges),memory_write(tier: stm|ltm),memory_write_episode,memory_reinforce,memory_edit(revise|promote|demote),memory_graph(edges),procedure(grounding|propose|outcome|sweep),task_state(read|write),memory_maintain(decay_sweep|pattern_check).full(default for the HTTP daemon;serve --tools fullorBACKANT_MEMORY_TOOLS=full) — core plus every pre-0.4 name unchanged (memory_write_stm,memory_recall_with_edges,task_state_read, …). Same handlers, so nothing that talks to the legacy surface breaks.
memory_recall returns {hits, count, note?} — an empty result carries a
"no memories yet — write one" note instead of a bare [].
For agents whose harness has no MCP surface, or that are told to use the CLI
directly, four verbs reach the same repo-scoped store serve opens for stdio.
The store is resolved from the git origin of the working directory, so run them
from inside the checkout the memory belongs to. Each exits non-zero with the
reason on stderr when the store cannot be reached.
backant-memory recall --cue "what you are about to re-derive" [--k 10] [--tier any|stm|ltm]
backant-memory reinforce --id <id> [--reason act-cite]
backant-memory write --tier stm|ltm --type <type> --content "<text>" --source <path_or_url> [--reason <why>]
backant-memory episode --situation "<what you faced>" --action "<what you did>" \
--expected success|failure --outcome success|failure|partial [--evidence "<what shows it>"]recall prints one JSON object per line with id, tier, type, age and
content, so it pipes into jq without a wrapper. write --tier ltm requires
--reason, the same rule memory_write enforces. reinforce's --reason is a
CATEGORY and not a note: act-cite (the default) and dream-cite raise
verdict_boost, anything else only touches last_reinforced, and weight is
capped at 1.0 so a freshly written row moves its citation counters rather than
its number.
backant-memory status # launchctl state + /healthz, one line
backant-memory doctor # every install check, pass/fail per item
backant-memory doctor --verify-restart # SIGKILL the daemon, prove launchd relaunches it
backant-memory usage --days 30 # adoption: sessions by entrypoint, hook/digest presence, memory calls per 1k turnsusage reads the Claude Code transcripts under ~/.claude/projects (read-only)
and is the number to watch: a store's size says nothing about whether agents
actually recall and write.
There is nothing to do. Reboot, log back in, and:
backant-memory statusshould already report service: running; http: healthy — launchd relaunches
the daemon at login automatically.
print-config emits ready-to-paste snippets. The default (generic) prints
both: the stdio entry (recommended, per-session repo-scoped — the same shape
Claude Code is registered with) and the streamable-HTTP + Authorization entry
for agents that only speak HTTP (which talk to the global store).
backant-memory print-config # both stdio + http (generic)
backant-memory print-config --client claude # stdio onlyAll settings are optional and read from the environment.
| Variable | Default | Notes |
|---|---|---|
BACKANT_MEMORY_HOME |
~/.claude/kairos |
Data home. Shared with backant-kairos on purpose. Honoured by the hooks' cold path too. |
BACKANT_MEMORY_PORT |
41414 |
HTTP port for the daemon. |
BACKANT_MEMORY_OLLAMA_URL |
http://127.0.0.1:11434 |
Local Ollama endpoint. Falls back to KAIROS_OLLAMA_URL. |
BACKANT_MEMORY_EMBEDDING_MODEL |
qwen3-embedding:0.6b |
Embedding model. Falls back to KAIROS_EMBEDDING_MODEL. |
BACKANT_MEMORY_DB |
(unset) | Read by serve only: pin a fixed store file, bypassing repo-scope resolution (stdio) and the global default (http). For tests and pinned single-store setups. |
Embeddings are always produced locally through Ollama — there are no remote embedding APIs, ever.
backant-memory uninstallReverses everything install did (boots out and removes the plist, and strips
only the content inside its own markers/keys). Your memories and the installed
skill directory are left in place.
The memory database lives under BACKANT_MEMORY_HOME and is shared with
backant-kairos (which still carries its own embedded copy of the memory
system). Because both read and write the same store, the schema is frozen
and checksum-pinned in both repositories until backant-kairos is refactored to
consume this package. See the design spec in the backant-kairos repo:
docs/superpowers/specs/2026-07-03-standalone-memory-mcp-design.md.
Elastic-2.0. See LICENSE.