Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -55,3 +55,7 @@ jobs:
run: tests/pool-settling-window.sh
- name: stuck queue guard
run: tests/stuck-queue-guard.sh
- name: pool labels
run: tests/pool-labels.sh
- name: pool rename
run: tests/pool-rename.sh
21 changes: 20 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ Break any of these and the tool stops being what it is.
```
bin/runpool the executable, its dispatcher, its help text, and RUNPOOL_VERSION
lib/common.sh config, logging, pool loading, launch agents, deregistration
lib/lifecycle.sh register, set-count, up, down, reregister, remove
lib/lifecycle.sh register, set-count, up, down, reregister, rename, remove
lib/apply.sh the pools file, and reconciling the machine to it
lib/scheduler.sh status, doctor, autoscale, sweep, clean, schedule
lib/notify.sh the optional notifier hook and what triggers it
Expand Down Expand Up @@ -59,6 +59,24 @@ assets/icon.svg the icon, source of truth; PNGs are rendered from it

**The consequence is that an agent already loaded is not necessarily an agent that behaves correctly.** A plist rewritten on disk changes nothing until the pool cycles. Anything depending on agent behaviour must therefore read the *loaded* environment with `launchctl print`, not the file. `_rp_agent_traps_signals` is the example, and `_rp_drain_pool` refuses per runner on the strength of it. The file on disk is what somebody intended; the loaded environment is what is true.

## The pool name is a label

**`POOL_LABELS` is the full list handed to `config.sh`, and the extras are derived back out of it.** There is deliberately no `POOL_EXTRA_LABELS`. A second field would be a frozen copy of `_rp_extra_labels`, and configs get edited by hand: the copy would go stale on the next hand edit and the following `apply` would silently revert it, which is the very defect `--labels` exists to fix.

- **`--labels` appends and cannot replace.** GitHub assigns `self-hosted`, the OS and the architecture to every self-hosted runner whatever it is told, so a replacing flag would promise something GitHub overrides. The pool name is in the base set too, because it is the routing contract.
- **An implicit label is refused, not dropped.** A token accepted and then normalised away would leave what the pools file declares and what the config holds permanently unequal, and `apply` would re-register the pool on every run.
- **Extras are compared sorted and stored in declared order.** Unlike `--watch`, which is compared as written: the costs are not symmetric, because a spurious watch difference is one file write and a spurious label difference stands the whole pool down.
- **`_rp_extra_labels` is tolerant and `_rp_valid_label` is strict.** The first runs against whatever a human left in a config and must only filter; the second guards the way in. That character class is what keeps a sourced config safe, `apply`'s `|` separator intact, and `rename`'s `find -name` free of globs.

## Renaming moves, and must deregister

**`rename` uses `mv`, unlike `migrate-storage`, which copies.** The copy there is because it crosses storage roots, where a partial copy is real and the old tree is the safety net. A rename keeps the same parent by construction, so `mv` is atomic and a second on-disk copy of runner credentials buys nothing.

- **It holds two locks, old then new, released in reverse.** From the moment the new config exists, `_rp_pool_names` returns it and autoscale would start it mid-rename. The second acquisition failing has to release the first.
- **`config.sh --replace` cannot help.** It replaces a registration of the same name, and the name is what changes, so GitHub keeps the old one: permanently offline, still carrying the old pool name as a label, and unreachable afterwards because `config.sh` overwrites the `.runner` holding its `agentId`. The old registrations are therefore deleted explicitly, before reconfiguring.
- **It iterates the runner directories that exist, not `1..POOL_COUNT`.** A count lowered by hand leaves higher-numbered runners on disk and registered; those are deregistered and deliberately not re-registered, which would be a permanent `miscount`.
- **`_rp_migrate_update_pool_conf` is not reusable here.** Its `END` clause adds `POOL_CACHE_DIR` when absent, and that absence is exactly how `_rp_load_pool` recognises a legacy pool. `rename` writes its own config instead, in one write, omitting the field when the pool did not have it.

## The stuck-queue guard subtracts, it does not suppress

**A queued run that never starts would otherwise wake a pool for ever**, every `RUNPOOL_IDLE_SECS` plus a tick, and nothing reports it because a pool that wakes and stands down is behaving as designed. `_rp_autoscale` therefore computes `queued > held` rather than deciding whether the pool is allowed to wake.
Expand All @@ -75,6 +93,7 @@ assets/icon.svg the icon, source of truth; PNGs are rendered from it

- **The holder writes its pid into the lock.** `_rp_set_count_locked` calls `_rp_up` at the end of its own work while still holding the lock, so a test for mere existence would deadlock every resize against itself. `_rp_resize_locked_by_other` is the predicate to use.
- **Staleness is measured from the lock's mtime**, and a drain refreshes it every poll. So an old lock means nobody is tending it, not that the work is slow. Anything that can hold the lock for a long time must call `_rp_resize_lock_touch`.
- **`rename` holds two, old then new, released in reverse.** The lock directory carries the pool name, so one lock cannot cover both. They cannot deadlock, because `_rp_resize_lock` refuses rather than blocks, and the second acquisition failing must release the first.
- **A dead holder is not an obstacle.** It left the directory behind, and the stale break in `_rp_resize_lock` is what clears it.

## Configuration precedence
Expand Down
5 changes: 4 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,7 +56,7 @@ The first job after a quiet spell waits about a minute for its pool to come up.

| Command | |
|---|---|
| `register <pool> --repo OWNER/REPO\|--org ORG [--count N] [--watch OWNER/REPO,...] [--allow-public]` | Create a pool and configure its runners |
| `register <pool> --repo OWNER/REPO\|--org ORG [--count N] [--watch OWNER/REPO,...] [--labels LABEL,...] [--allow-public]` | Create a pool and configure its runners |
| `set-count <pool> N [--if-count M] [--drain]` | Change a pool's runner count. `--if-count` refuses unless it is currently M; `--drain` lets running jobs finish first |
| `apply [--dry-run] [--file PATH]` | Reconcile the machine to a file describing its pools |
| `up` / `down <pool> [--drain\|--force]` | Bring a pool online, or stand it down. `--drain` waits for running jobs; `--force` ends them |
Expand All @@ -67,6 +67,7 @@ The first job after a quiet spell waits about a minute for its pool to come up.
| `stats [--queue] [--days N \| --all]` | What jobs cost, from recorded telemetry. `--queue` adds the wait before each job started, over the last 7 days unless widened |
| `pause [pool]` / `resume [pool]` | Global kill switch, or persistent per-pool pause |
| `reregister <pool>` | Recreate GitHub registrations, keeping the local install |
| `rename <old> <new> [--drain]` | Rename a pool, locally and at GitHub |
| `rewrite-agents` | Regenerate the launch agents after changing hook settings |
| `remove <pool>` | Deregister and delete a pool |
| `clean [pool]` | Prune work directories, temp, diagnostics, old binaries, caches |
Expand All @@ -83,6 +84,8 @@ Three commands earn a note beyond the table:
- **The pools file is intent; the running pool is state.** `set-count` changes the pool and deliberately does not write the file, so the two disagree after any resize. That is the normal condition between them rather than a fault: the file records the shape you want a machine to have and is what you copy between machines, while the pool records what is running right now. `apply` is where they are reconciled, and it resolves the difference in the file's favour, so `apply --dry-run` first is not a formality. A `count 3 -> 4` line in that plan is the drift, and applying it would undo a deliberate resize.
- **`set-count` is absolute, so a caller that reads a count and acts on it later needs `--if-count`.** A pool changed in between turns a growth into a shrink, and shrinking deregisters runners. `--if-count M` refuses unless the pool is still at M, and one resize per pool runs at a time so two callers cannot interleave. A runner deregistered locally that GitHub still holds is reported as a failure, not logged and passed over: a stale registration attracts jobs that then queue forever.
- **`status --json --local` skips the GitHub query**, reporting those fields as `null`. The root `paused` field is the global kill switch; every pool also carries its own additive `paused` field. Anything refreshing on a timer should use `--local`, since one API call per pool per minute is thousands a day and makes a passive readout fail whenever the network does.
- **The pool's name is one of its runners' labels, so renaming changes routing.** Every runner carries `self-hosted`, the machine's OS and architecture, the pool's name, and anything `--labels` adds. `rename` moves the pool's directories, config, launch agents and state, then re-registers every runner under the new name, which means `runs-on: [self-hosted, <old>]` stops matching and has to be updated. It deletes the old GitHub registrations rather than replacing them, because `--replace` only covers a name collision and the name is what is changing; a registration left behind would be permanently offline while still advertising the old label. `rename` does not touch the pools file, so update that too.
- **A label change re-registers the whole pool.** Labels live on GitHub's registration, not in a file GitHub reads, so `apply` applies one by standing the pool down and re-registering every runner. That is far heavier than a count or watch-list change and the plan says so. Absent `--labels` means no extra labels, the same way absent `--count` would mean the default: a pool whose config was edited by hand to add a label needs that label declared in the file, or the next `apply` removes it.
- **`doctor` answers "why is nothing picking this up" in one command.** It checks `gh` and its authentication, that GitHub still holds the registrations, that the launch agents exist, and then disk headroom, config permissions and the organisation's runner-group setting. Each failure comes with what to do about it, and it exits non-zero when something is actually wrong. It repairs nothing, so it is safe at any moment including mid-job.

## Describing a machine's pools
Expand Down
29 changes: 28 additions & 1 deletion bin/runpool
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@ export RUNPOOL_INVOKED
# The released version, and the only place it is written. The Homebrew formula
# builds from a git tag, so a tag without a matching bump here ships a binary
# that misreports itself.
RUNPOOL_VERSION="0.11.0"
RUNPOOL_VERSION="0.12.0"

# shellcheck source=lib/common.sh
. "${RUNPOOL_ROOT}/lib/common.sh"
Expand Down Expand Up @@ -83,6 +83,7 @@ Commands:
pools List registered pools
stats [options] Show recorded job durations and queue times
reregister <pool> Recreate a pool's GitHub registrations
rename <old> <new> Rename a pool, locally and at GitHub
remove <pool> Deregister and delete a pool
tick Run autoscaling and pool health checks
autoscale Bring pools up when work is queued
Expand Down Expand Up @@ -110,7 +111,32 @@ Options:
--org ORG Register an organisation-scoped pool
--count <n> Set the runner count (default: 2)
--watch OWNER/REPO,... Repositories an organisation pool watches for work
--labels LABEL,... Extra runner labels, on top of self-hosted, macOS, ARM64
and the pool's own name. Letters, digits, dot,
underscore and hyphen; several separated by commas
--allow-public Allow a public repository (repository scope only)
HELP
;;
rename)
cat <<'HELP'
Usage: runpool rename <old> <new> [options]

Rename a pool. Moves its directories, config, launch agents and state, then
re-registers every runner with GitHub under the new name.

Options:
--drain Let running jobs finish first, then rename
--timeout <seconds> How long --drain waits

The pool name is one of the runner's labels, so 'runs-on: [self-hosted, <old>]'
stops matching and has to be changed to the new name. A workflow asking only
for 'self-hosted' is unaffected.

The old GitHub registrations are deleted rather than replaced, because
--replace only covers a name collision and the name is what is changing.

The pools file is not touched. Update it too, or the next 'runpool apply' will
create the old pool again.
HELP
;;
set-count)
Expand Down Expand Up @@ -234,6 +260,7 @@ case "${cmd}" in
set-count) _rp_set_count "$@" ;;
apply) _rp_apply "$@" ;;
reregister) _rp_reregister "$@" ;;
rename) _rp_rename "$@" ;;
up) _rp_up "$@" ;;
down) _rp_down "$@" ;;
up-all) _rp_up_all ;;
Expand Down
Loading