Skip to content

Daemon skill-store: a broken or half-installed skill-store/<hash> makes every thread start fail (ENOENT …/.last-used / EIO … rename) until it's deleted by hand #4855

Description

@WyrdWerk

Environment. bb-app 0.45.0 server and host daemon (daemon installed by the server on a remote
Linux machine, Ubuntu 24.04), claude-code and pi providers. Related but different: #4589 (HTTP 500
on skill-tree fetch, closed).

Symptom. thread.start fails within seconds with one of:

  • ENOENT: no such file or directory, open '…/runtime/skill-store/<hash>/.last-used'
  • EIO: i/o error, rename '…/runtime/skill-store/.tmp-<hash>-…' -> '…/runtime/skill-store/<hash>'

The daemon logs Failed to pull required injected skill tree … sourceType: data-dir (name
acp-provider) on every start, and it doesn't recover by itself. In the ENOENT case
runtime/skill-store/<hash>/ existed but was empty (healthy entries have .complete,
.last-used, content/). Deleting that empty dir (and any .tmp-* leftovers) fixed it, and the
next start pulled the tree normally.

Trigger. On our machines the disk is restored lazily after start/resume, so file operations,
including renames, can fail transiently with EIO for a few minutes. The trigger is environmental,
but getting stuck permanently is in the daemon.

What the 0.45.0 daemon does (daemon-bundle.mjs, the skill-tree install path):

  • The cached path checks <hash>/.complete and <hash>/content; on ENOENT it pulls again. ✔
  • The install builds .tmp-<hash>-<pid>-<time>-<rand>/ (content, .last-used, .complete), then
    rename(tmp, <hash>):
    • on EEXIST / ENOTEMPTY it deletes the tmp tree and continues as if installed, writing
      <hash>/.last-used and returning <hash>/content without re-checking <hash>/.complete,
      so an existing broken <hash> is never replaced;
    • on any other error (e.g. EIO) it deletes the tmp tree and throws; the turn fails, with no
      retry until the next start.

Exactly how the empty <hash> dir appeared is inferred, not proven (a partial rename, or a
half-restored dir, then trusted by the EEXIST branch).

Suggested fixes.

  1. In the EEXIST/ENOTEMPTY branch, re-check <hash>/.complete + content/; if missing, remove the
    broken <hash> and retry the rename once.
  2. Treat a <hash> dir without .complete as absent everywhere (remove before pulling).
  3. Retry transient rename errors (EIO, EBUSY) a few times with short backoff, and sweep stale
    .tmp-* entries at daemon start.

Workaround (what our machine-provider plugin does): before enrolling or resuming a machine,
delete empty runtime/skill-store/<hash> dirs and .tmp-* leftovers; repeat for ~15 min after
enrollment (only .tmp-* older than 10 min).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    confirmed-reproBug reproduced again from a clean trusted checkout; see linked reporthostHost daemon, process lifecycle, memory, event loop

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions