Skip to content

Repository files navigation

Bench

A bench of Claude Code specialists, and one page to decide from.

Bench runs Claude Code as long-lived processes — specialists — each in its own git worktree, and surfaces their work as decisions rather than transcripts. You do not read the scrollback. You read one page and answer one question.

The cockpit: a specialist waiting on a decision

Why

Running an agent on a real task means one of two bad options: watch a terminal scroll for twenty minutes, or come back later and reconstruct what happened from a diff. Neither scales past one agent.

Bench takes the position that the interesting unit is not the message, it is the turn — and that a turn which produced work owes you a page you can decide from. A specialist writes a report when a decision needs you, when finished work needs understanding, when a spec needs approving, or when it is stuck. The rest of the time it just answers.

It is the same Claude Code you use in a terminal: every skill, every MCP server, subagents, web search. Bench supervises it; it does not replace it.

What it does

  • Specialists. One long-lived claude -p process each, by default in its own git worktree on its own branch, with node_modules and .env symlinked from your checkout so it can build and test what it writes without an install of its own. Untick Start in a worktree and it works directly in your checkout instead, on the branch you already have open.
  • Nothing is installed, and nothing is copied. A worktree borrows the dependencies your checkout already has, so provisioning takes milliseconds rather than the twenty seconds an install cost. The flip side is worth knowing: node_modules is a symlink to yours, so a specialist that runs pnpm install is installing into your checkout, and every other worktree reads through the same link. Nothing stops it — a specialist has a full shell, and the blanket denial that used to be here was removed because an unexplained refusal sent agents round it rather than stopping them. Ask for a dependency to be added rather than adding it, and if one appears in a diff you did not expect, that is why.
  • Decisions, not transcripts. Reports render as pages with numbered options. Press 1–n, Enter. The answer goes back into the live session.
  • Intake. A specialist can ask everything it needs at once, with its own picks pre-filled, so only the questions it genuinely cannot guess block it.
  • Progress you can read. A live trail derived from tool calls — Bash pnpm test, Edit src/registry.ts — beside the specialist's own checklist.
  • Watch it write, if you want to. The same tool calls, kept whole, go down /events as edit notices. A VS Code extension in editor/vscode/ opens each file a specialist writes — including one inside its worktree, which lives under the folder you already have open. It takes focus every time, so it is made for a second monitor; the status bar pauses it.
  • And read what they wrote. The extension's sidebar lists every specialist on the project with the files it has changed since its branch started, committed or not, each opening as a diff. It badges how many are waiting on a decision from you.
  • Any model, not only Claude's. Anthropic's aliases go straight to Anthropic on the login you already have. Everything else goes through an OpenRouter key you supply, and is billed there. The picker offers the models that can actually run a specialist — the ones that support tool use, which is most but not all of what OpenRouter carries — searchable, with what each holds and what it costs. Pick one when you make a specialist or change it after, from the model name at the foot of the composer.
  • They outlive the daemon. Restart Bench and the roster comes back. Nothing respawns until you prompt it, and then it resumes with its memory intact.
  • Gates. A commit carrying AI attribution is denied at PreToolUse. A specialist may not push a branch.

The roster puts whoever is waiting on you at the top of their project. If that is not the order you want, drag a row by the grip on the right — or focus the grip and use ↑ ↓ — and that group keeps the arrangement you gave it. It is remembered in the browser you arranged it in, and nowhere else.

The roster, grouped by project

Running it

Requires Node 22+, pnpm, git, and the claude CLI already authenticated.

pnpm install
pnpm build
pnpm start

It prints a localhost URL with a token. Open it, and bookmark it — the token is kept in ~/.bench/token (mode 0600) so the link keeps working across restarts. Delete that file to rotate it. The daemon binds to 127.0.0.1 only and every API route requires the token.

By default it looks for git repositories under /var/www. Point it elsewhere:

BENCH_PROJECTS_ROOT=~/code pnpm start
Variable Default What it does
BENCH_PROJECTS_ROOT /var/www Where to look for projects
BENCH_PORT 7420 Cockpit port
BENCH_HOME ~/.bench Where the specialist index and token are kept
BENCH_TOKEN generated Override the cockpit token
BENCH_LAN unset 1 binds every interface, not just loopback
BENCH_HOST 127.0.0.1 Bind one interface by name instead
BENCH_HEADROOM_BIN PATH lookup Where the headroom binary is, when it is not on PATH
BENCH_HEADROOM_PORT 8787 The loopback port the Headroom proxy runs on

Opening it from another device

pnpm start:lan

It prints every address the cockpit can be reached at, and says once that it is no longer only on this machine.

Weigh that before you use it. The token is the whole of the authentication, it travels in the URL over plain HTTP, and a specialist has a full shell — so anyone on that network holding the token can run anything on this machine. On a home network with a 48-character secret that is a reasonable trade; on a café network it is not a trade at all.

Settings holds the house rules every specialist is given, and the address this tab is talking to — point it at another machine's daemon and the same page loads from there.

Keys live in your profile

Your profile — the dialog behind the profile button at the top of the roster, once you have signed in with Google — is where API keys are kept. An Anthropic key saved there is synced to every Bench you sign in to, so a key written down once follows you to each machine rather than being pasted in per daemon. Either a console API key, which bills the API, or a token from claude setup-token, which bills the subscription it was minted from.

Each key is checked against the API before it is spent — the CLI retries a bad one ten times before it gives up, so a typo is worth catching there — and the list says which one is in use and when it was last checked. When the key in use reports a usage limit, the next turn moves to the next available key on the list. With no usable key, specialists run on this machine's own Claude login.

An OpenRouter key lives there too, on its own list — see the next section for what it is for.

Running a specialist on something other than Claude

Claude Code speaks one protocol and OpenRouter serves it, so pointing a specialist at Gemini or GPT or Llama is three environment variables on the child process — there is nothing to install and no second process that can be down. Save an OpenRouter key in your profile — or from the picker itself, which offers to take you there — and the picker fills in. Search it by name or id; each row says what the model holds and what a million tokens of its output costs, because that spend is yours rather than a subscription's. Without a key the list still shows, disabled: what you could run is worth more than a list that quietly omits most of it.

Models that cannot use tools are left out. That is not most of them, but it is not a small number either, and none of them could have run a specialist — without tool use a specialist cannot read a file, edit one or run a command. Google's image and music models are in that group, and they used to sort to the top of the picker.

Anthropic's own models deliberately do not go this way. They go direct, on your own login or key, because that is what bills the subscription you are already paying for — routing them through OpenRouter would quietly move that spend somewhere else.

The model a specialist runs on is not fixed at creation. Change it from the model name at the foot of the composer; it takes effect on the next prompt, which restarts the specialist on the new model and picks the conversation up where it left off. A model whose provider you have no key for is refused while you are still looking at the picker, rather than two minutes later as a turn that hangs and dies.

What a turn will be spent from

A small mark at the end of the composer opens what the account behind this specialist has spent — and which account that is follows the model beside it, because they are the same question.

For an Anthropic model it is the subscription: a bar per window, five-hour and seven-day, with what is left in each, and the mark is those windows in miniature. For an OpenRouter model it is that key's own spend against its ceiling, drawn as a ring rather than columns — money is spent, not refilled, and a key with no ceiling gets no invented percentage. Either way it colours itself amber past three-quarters and red past ninety percent, so it warns you without being opened.

The Anthropic mark needs an OAuth credential to ask with — its own setup token, or this machine's login. A console API key is billed rather than rationed and has no windows to report, so the mark stays away rather than showing an empty panel.

Settings: house rules, and the address this tab is talking to

Nothing hot-reloads yet: pnpm build before pnpm start, and run it from the repository root.

Restarting it

bench restart --build

Every change under src/daemon needs the daemon started again, and the usual way to do that — find the terminal, Ctrl-C, pnpm start — is not available to a specialist, which is exactly who tends to have just made the change.

Stopping the daemon stops every specialist, so this cannot be a straight kill and start: whoever typed it is about to be stopped by it. It hands the work to a detached process instead, which waits until no turn is running before it stops anything. So a specialist can restart the daemon it is running under, its turn finishes and is written down, and the roster comes back with everyone still on it. Nothing respawns until you prompt it, the same as any other restart.

With --build it builds first and leaves the running daemon alone if the build fails — a daemon that will not start is worse than the one you had.

Two things change after a restart run this way. The daemon has no terminal, so its output goes to ~/.bench/daemon.log, and the restart's own progress to ~/.bench/restart.log. And the terminal you originally started it in is now free: it printed bench: stopping. and exited.

Compressing prompts with Headroom

Headroom is a local proxy that compresses what each turn sends before it reaches Anthropic. Install it once:

uv tool install --python 3.12 "headroom-ai[proxy]"

If the binary is there, the daemon starts headroom proxy on 127.0.0.1:8787 itself — or reuses one you are already running on that port, and never kills it. Specialists that talk to Anthropic directly are pointed at it through ANTHROPIC_BASE_URL on their next start; ones already running keep what they were spawned with. The proxy is run offline and with its telemetry beacon off, because a tool whose job is shrinking what leaves the machine should not be adding to it.

There is a toggle in Settings, on by default since it does nothing without the binary. Switching it off stops handing the proxy out — it keeps running, and specialists pointed at it stay pointed until they respawn. Specialists answered by OpenRouter are not routed through it: they already go through Bench's own translation proxy, and stacking a second one in that chain is unproven. headroom savings shows what it is actually removing.

The bench command

Installed alongside the daemon, and mostly for the specialists rather than for you: a specialist that has been asked to spec something and hand it to an implementer needs a way to open a tab, and inventing an API call is exactly what an agent should not be doing.

bench ls                         who is on this project, on what model, doing what
bench new <label> [--as <role>]  open a tab; it waits until it is told what for
bench tell <label> "<text>"      give one its next turn
bench close <label>              done with a sub-agent you opened — shut it down
bench restart [--build]          stop the daemon and start it again
bench remote off                 stop being reachable from your other devices

A specialist may only open, tell and close tabs on its own project, and may only close one it opened itself — not the developer's, and not its own. That check is in the command rather than the daemon, because the token is shared and the daemon cannot otherwise tell who is asking.

It reads the token itself rather than taking one on the command line. curl "...?token=$(cat ~/.bench/token)" puts the secret into ps output for every process on the machine to read; this does not.

bench restart is the one of these you are as likely to type as they are — see Restarting it for what it does and where the output goes afterwards.

How it works

browser ──ws/http──> daemon ──stdin/stdout(stream-json)──> claude -p
                       │                                      │
                       ├── git worktree per specialist ───────┘
                       └── .bench/reports/<id>/  reports, replies, threads

A turn ends with a result event and the process blocks on stdin, so the turn is the unit of control and no separate "needs input" protocol exists. Prompts sent while a turn is running are held by the daemon, not written to stdin, so they get a turn of their own instead of being swallowed by the one in flight.

Reports and replies are agent-authored HTML, rendered in a sandboxed frame under a strict CSP: no network, no scripts, no external anything.

Status

docs/STATUS.md is the honest account — what is built and proven against a real CLI, what is built but unproven, what is broken, what was deliberately left out, and the bugs found by using it. It is kept current because it goes stale the moment the code moves.

Bench is early. It is used to build itself, which is where most of its bugs have come from.

Contributing

See CONTRIBUTING.md. Two things worth knowing up front: tests must run against the real CLI before a claim is called proven, and commits never carry AI attribution — there is a gate that enforces it.

Licence

MIT — see LICENSE.

About

AI coding that actually makes you move with superhuman speed.

Topics

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages