seeks is a control plane, not a sandbox. Read the coverage table in the README before relying on it — it is deliberately explicit about the boundary.
In short:
- Edits are deterministically enforced. Denylist (a floor a loop can extend but not narrow), worktree confinement, hook-owned loop state, seeks' own plugin code, L1 report-only, wrap-up window. This is the only tier that holds by construction.
doneis released only by the Stop hook, after it runs the loop's stored done-conditions itself and each exits as expected.seeks certify(the verifier's sign-off) andseeks oracle-ackare advice: the maker can call them too. If a pre-existing oracle file (tests, build manifests, runner configs, CI) was modified or deleted, or a config-type one was added, a green check releases needs-human, not done, unless the loop opted intooracle_modified_policy: "ack". The stored conditions and the gate's verification cache are loop state, so this holds only as far as the loop-state protection below does, and the checks run in the maker's worktree (gitignored dependency directories included).- The seeks CLI won't move a live loop's brakes without the user.
status-setnever writesarmed/done/verifier_certified, and freezes the budget, sweep, oracle, condition (incl.condition_timeout_sec) and policy keys and the stuck guard's counters while the loop is live.start/stop/reset-fires/budget-set/start-clock/base-record/re-init/gcon a live loop need the one-shot grant that only a user-typed/seeks:start|stop|deletemints, for the loop it names. A non-interactive Claude Code (claude -p) never mints one, and starting Claude Code from inside a loop is denied, at the Bash tier. - Everything judged from a Bash command string is best-effort — including the budget.
git push/merge/rebaseand any touch ofstatus.jsonare parsed for, not pattern-matched, throughcdtracking,..collapsing, brace expansion,*/?/[…]glob resolution,eval/sh -crecursion and interpreter-payload scanning. It is thorough and it is not a proof. The README lists what still gets through, and each of those is a passing test asserting allow. - Everything else Bash can do is best-effort by default.
SEEKS_STRICT_BASH=1turns Bash into a deny-by-default allowlist, which is much stronger — but it is still an allowlist, not a sandbox:node -eis on it, and a shell is Turing-complete. - Reads are not policed at all. The model can read
.envand your secrets. - Runtime-assembled paths and encoded payloads are explicitly out of scope. A name built by
$(…), abase64 -d | sh, or a write performed inside a script the hook only sees the filename of are documented non-goals — not oversights.
If the goal or the codebase is untrusted, run the loop in a container (seeks run --container: worktree, .git and .seeks mounted read-write, the plugin read-only, no host HOME, credentials by env name). That is the only guarantee that holds by construction rather than by policy, and it is a guarantee about the rest of your machine: loop state (including the stored done-conditions) sits in the writable .seeks mount under the same best-effort policy.
Please report privately via GitHub's private vulnerability reporting rather than a public issue.
Include:
- what the guardrail claims (quote the README or
SKILL.mdline — an overclaim in the docs is a valid report on its own), - the exact command or tool input that gets past it,
- the
rulefrom/seeks:why <name> --denied, if one fired, /seeks:exportoutput if you can share it.
I'll acknowledge within a week. Since this is a solo project, expect a fix or a documented scope change rather than a formal advisory timeline.
In scope — anything that lets a loop do what the README says it cannot: reach status.json / hook-state.json / decisions.jsonl / control-grant.json or seeks' own plugin code, push/merge/rebase, edit a denylisted path or escape the worktree via the edit tools, defeat the iteration or wall-clock cap, get the Stop gate to release done while a stored done-condition fails, or move a live loop's brakes through the seeks CLI without the user's grant. Also in scope: any claim in the docs that the code does not enforce.
Out of scope — the documented gaps above (unpoliced Bash without strict mode, unpoliced reads, a runtime-assembled path, an encoded payload, a write inside a script the hook only sees the name of, a cd carried over from an earlier Bash call, node -e under strict mode). Also known and stated: the maker can game the code under test itself, a helper script outside the oracle globs, or the installed tools in a gitignored dependency directory (node_modules/, .venv/), and it reports its own discovery sweeps (sweep-tick), so the sweep bar is a thoroughness heuristic, not a guarantee. Those are known, stated, and pinned as passing allow tests. If you can show one is worse than documented — or find a plainly-spelled command that reaches loop state — that is in scope and worth reporting.