A bench of Claude Code specialists, and one page to decide from.
Bench runs Claude Code as long-lived processes — specialists — each in its own git worktree, and surfaces their work as decisions rather than transcripts. You do not read the scrollback. You read one page and answer one question.
Running an agent on a real task means one of two bad options: watch a terminal scroll for twenty minutes, or come back later and reconstruct what happened from a diff. Neither scales past one agent.
Bench takes the position that the interesting unit is not the message, it is the turn — and that a turn which produced work owes you a page you can decide from. A specialist writes a report when a decision needs you, when finished work needs understanding, when a spec needs approving, or when it is stuck. The rest of the time it just answers.
It is the same Claude Code you use in a terminal: every skill, every MCP server, subagents, web search. Bench supervises it; it does not replace it.
- Specialists. One long-lived
claude -pprocess each, by default in its own git worktree on its own branch, withnode_modulesand.envsymlinked from your checkout so it can build and test what it writes without an install of its own. Untick Start in a worktree and it works directly in your checkout instead, on the branch you already have open. - Nothing is installed, and nothing is copied. A worktree borrows the
dependencies your checkout already has, so provisioning takes milliseconds
rather than the twenty seconds an install cost. The flip side is worth
knowing:
node_modulesis a symlink to yours, so a specialist that runspnpm installis installing into your checkout, and every other worktree reads through the same link. Nothing stops it — a specialist has a full shell, and the blanket denial that used to be here was removed because an unexplained refusal sent agents round it rather than stopping them. Ask for a dependency to be added rather than adding it, and if one appears in a diff you did not expect, that is why. - Decisions, not transcripts. Reports render as pages with numbered options.
Press
1–n,Enter. The answer goes back into the live session. - Intake. A specialist can ask everything it needs at once, with its own picks pre-filled, so only the questions it genuinely cannot guess block it.
- Progress you can read. A live trail derived from tool calls —
Bash pnpm test,Edit src/registry.ts— beside the specialist's own checklist. - Watch it write, if you want to. The same tool calls, kept whole, go down
/eventsas edit notices. A VS Code extension ineditor/vscode/opens each file a specialist writes — including one inside its worktree, which lives under the folder you already have open. It takes focus every time, so it is made for a second monitor; the status bar pauses it. - And read what they wrote. The extension's sidebar lists every specialist on the project with the files it has changed since its branch started, committed or not, each opening as a diff. It badges how many are waiting on a decision from you.
- Any model, not only Claude's. Anthropic's aliases go straight to Anthropic on the login you already have. Everything else goes through an OpenRouter key you supply, and is billed there. The picker offers the models that can actually run a specialist — the ones that support tool use, which is most but not all of what OpenRouter carries — searchable, with what each holds and what it costs. Pick one when you make a specialist or change it after, from the model name at the foot of the composer.
- They outlive the daemon. Restart Bench and the roster comes back. Nothing respawns until you prompt it, and then it resumes with its memory intact.
- Gates. A commit carrying AI attribution is denied at
PreToolUse. A specialist may not push a branch.
The roster puts whoever is waiting on you at the top of their project. If that
is not the order you want, drag a row by the grip on the right — or focus the
grip and use ↑ ↓ — and that group keeps the arrangement you gave it. It is
remembered in the browser you arranged it in, and nowhere else.
Requires Node 22+, pnpm, git, and the claude CLI already authenticated.
pnpm install
pnpm build
pnpm startIt prints a localhost URL with a token. Open it, and bookmark it — the token
is kept in ~/.bench/token (mode 0600) so the link keeps working across
restarts. Delete that file to rotate it. The daemon binds to 127.0.0.1 only
and every API route requires the token.
By default it looks for git repositories under /var/www. Point it elsewhere:
BENCH_PROJECTS_ROOT=~/code pnpm start| Variable | Default | What it does |
|---|---|---|
BENCH_PROJECTS_ROOT |
/var/www |
Where to look for projects |
BENCH_PORT |
7420 |
Cockpit port |
BENCH_HOME |
~/.bench |
Where the specialist index and token are kept |
BENCH_TOKEN |
generated | Override the cockpit token |
BENCH_LAN |
unset | 1 binds every interface, not just loopback |
BENCH_HOST |
127.0.0.1 |
Bind one interface by name instead |
BENCH_HEADROOM_BIN |
PATH lookup | Where the headroom binary is, when it is not on PATH |
BENCH_HEADROOM_PORT |
8787 |
The loopback port the Headroom proxy runs on |
pnpm start:lanIt prints every address the cockpit can be reached at, and says once that it is no longer only on this machine.
Weigh that before you use it. The token is the whole of the authentication, it travels in the URL over plain HTTP, and a specialist has a full shell — so anyone on that network holding the token can run anything on this machine. On a home network with a 48-character secret that is a reasonable trade; on a café network it is not a trade at all.
Settings holds the house rules every specialist is given, and the address this tab is talking to — point it at another machine's daemon and the same page loads from there.
Your profile — the dialog behind the profile button at the top of the
roster, once you have signed in with Google — is where API keys are kept.
An Anthropic key saved there is synced to every Bench you sign in to, so a
key written down once follows you to each machine rather than being pasted
in per daemon. Either a console API key, which bills the API, or a token
from claude setup-token, which bills the subscription it was minted from.
Each key is checked against the API before it is spent — the CLI retries a bad one ten times before it gives up, so a typo is worth catching there — and the list says which one is in use and when it was last checked. When the key in use reports a usage limit, the next turn moves to the next available key on the list. With no usable key, specialists run on this machine's own Claude login.
An OpenRouter key lives there too, on its own list — see the next section for what it is for.
Claude Code speaks one protocol and OpenRouter serves it, so pointing a specialist at Gemini or GPT or Llama is three environment variables on the child process — there is nothing to install and no second process that can be down. Save an OpenRouter key in your profile — or from the picker itself, which offers to take you there — and the picker fills in. Search it by name or id; each row says what the model holds and what a million tokens of its output costs, because that spend is yours rather than a subscription's. Without a key the list still shows, disabled: what you could run is worth more than a list that quietly omits most of it.
Models that cannot use tools are left out. That is not most of them, but it is not a small number either, and none of them could have run a specialist — without tool use a specialist cannot read a file, edit one or run a command. Google's image and music models are in that group, and they used to sort to the top of the picker.
Anthropic's own models deliberately do not go this way. They go direct, on your own login or key, because that is what bills the subscription you are already paying for — routing them through OpenRouter would quietly move that spend somewhere else.
The model a specialist runs on is not fixed at creation. Change it from the model name at the foot of the composer; it takes effect on the next prompt, which restarts the specialist on the new model and picks the conversation up where it left off. A model whose provider you have no key for is refused while you are still looking at the picker, rather than two minutes later as a turn that hangs and dies.
A small mark at the end of the composer opens what the account behind this specialist has spent — and which account that is follows the model beside it, because they are the same question.
For an Anthropic model it is the subscription: a bar per window, five-hour and seven-day, with what is left in each, and the mark is those windows in miniature. For an OpenRouter model it is that key's own spend against its ceiling, drawn as a ring rather than columns — money is spent, not refilled, and a key with no ceiling gets no invented percentage. Either way it colours itself amber past three-quarters and red past ninety percent, so it warns you without being opened.
The Anthropic mark needs an OAuth credential to ask with — its own setup token, or this machine's login. A console API key is billed rather than rationed and has no windows to report, so the mark stays away rather than showing an empty panel.
Nothing hot-reloads yet: pnpm build before pnpm start, and run it from the
repository root.
bench restart --buildEvery change under src/daemon needs the daemon started again, and the usual
way to do that — find the terminal, Ctrl-C, pnpm start — is not available to
a specialist, which is exactly who tends to have just made the change.
Stopping the daemon stops every specialist, so this cannot be a straight
kill and start: whoever typed it is about to be stopped by it. It hands
the work to a detached process instead, which waits until no turn is running
before it stops anything. So a specialist can restart the daemon it is running
under, its turn finishes and is written down, and the roster comes back with
everyone still on it. Nothing respawns until you prompt it, the same as any
other restart.
With --build it builds first and leaves the running daemon alone if the
build fails — a daemon that will not start is worse than the one you had.
Two things change after a restart run this way. The daemon has no terminal, so
its output goes to ~/.bench/daemon.log, and the restart's own progress to
~/.bench/restart.log. And the terminal you originally started it in is now
free: it printed bench: stopping. and exited.
Headroom is a local proxy that compresses what each turn sends before it reaches Anthropic. Install it once:
uv tool install --python 3.12 "headroom-ai[proxy]"If the binary is there, the daemon starts headroom proxy on 127.0.0.1:8787
itself — or reuses one you are already running on that port, and never kills
it. Specialists that talk to Anthropic directly are pointed at it through
ANTHROPIC_BASE_URL on their next start; ones already running keep what they
were spawned with. The proxy is run offline and with its telemetry beacon off,
because a tool whose job is shrinking what leaves the machine should not be
adding to it.
There is a toggle in Settings, on by default since it does nothing without
the binary. Switching it off stops handing the proxy out — it keeps running,
and specialists pointed at it stay pointed until they respawn. Specialists
answered by OpenRouter are not routed through it: they already go through
Bench's own translation proxy, and stacking a second one in that chain is
unproven. headroom savings shows what it is actually removing.
Installed alongside the daemon, and mostly for the specialists rather than for you: a specialist that has been asked to spec something and hand it to an implementer needs a way to open a tab, and inventing an API call is exactly what an agent should not be doing.
bench ls who is on this project, on what model, doing what
bench new <label> [--as <role>] open a tab; it waits until it is told what for
bench tell <label> "<text>" give one its next turn
bench close <label> done with a sub-agent you opened — shut it down
bench restart [--build] stop the daemon and start it again
bench remote off stop being reachable from your other devices
A specialist may only open, tell and close tabs on its own project, and may only close one it opened itself — not the developer's, and not its own. That check is in the command rather than the daemon, because the token is shared and the daemon cannot otherwise tell who is asking.
It reads the token itself rather than taking one on the command line. curl "...?token=$(cat ~/.bench/token)" puts the secret into ps output for every
process on the machine to read; this does not.
bench restart is the one of these you are as likely to type as they are —
see Restarting it for what it does and where the output
goes afterwards.
browser ──ws/http──> daemon ──stdin/stdout(stream-json)──> claude -p
│ │
├── git worktree per specialist ───────┘
└── .bench/reports/<id>/ reports, replies, threads
A turn ends with a result event and the process blocks on stdin, so the turn
is the unit of control and no separate "needs input" protocol exists. Prompts
sent while a turn is running are held by the daemon, not written to stdin, so
they get a turn of their own instead of being swallowed by the one in flight.
Reports and replies are agent-authored HTML, rendered in a sandboxed frame under a strict CSP: no network, no scripts, no external anything.
docs/STATUS.md is the honest account — what is built and proven against a real CLI, what is built but unproven, what is broken, what was deliberately left out, and the bugs found by using it. It is kept current because it goes stale the moment the code moves.
Bench is early. It is used to build itself, which is where most of its bugs have come from.
See CONTRIBUTING.md. Two things worth knowing up front: tests must run against the real CLI before a claim is called proven, and commits never carry AI attribution — there is a gate that enforces it.
MIT — see LICENSE.


