Environment. bb-app 0.45.0 server and host daemon (daemon installed by the server on a remote
Linux machine, Ubuntu 24.04), claude-code and pi providers. Related but different: #4589 (HTTP 500
on skill-tree fetch, closed).
Symptom. thread.start fails within seconds with one of:
ENOENT: no such file or directory, open '…/runtime/skill-store/<hash>/.last-used'
EIO: i/o error, rename '…/runtime/skill-store/.tmp-<hash>-…' -> '…/runtime/skill-store/<hash>'
The daemon logs Failed to pull required injected skill tree … sourceType: data-dir (name
acp-provider) on every start, and it doesn't recover by itself. In the ENOENT case
runtime/skill-store/<hash>/ existed but was empty (healthy entries have .complete,
.last-used, content/). Deleting that empty dir (and any .tmp-* leftovers) fixed it, and the
next start pulled the tree normally.
Trigger. On our machines the disk is restored lazily after start/resume, so file operations,
including renames, can fail transiently with EIO for a few minutes. The trigger is environmental,
but getting stuck permanently is in the daemon.
What the 0.45.0 daemon does (daemon-bundle.mjs, the skill-tree install path):
- The cached path checks
<hash>/.complete and <hash>/content; on ENOENT it pulls again. ✔
- The install builds
.tmp-<hash>-<pid>-<time>-<rand>/ (content, .last-used, .complete), then
rename(tmp, <hash>):
- on EEXIST / ENOTEMPTY it deletes the tmp tree and continues as if installed, writing
<hash>/.last-used and returning <hash>/content without re-checking <hash>/.complete,
so an existing broken <hash> is never replaced;
- on any other error (e.g. EIO) it deletes the tmp tree and throws; the turn fails, with no
retry until the next start.
Exactly how the empty <hash> dir appeared is inferred, not proven (a partial rename, or a
half-restored dir, then trusted by the EEXIST branch).
Suggested fixes.
- In the EEXIST/ENOTEMPTY branch, re-check
<hash>/.complete + content/; if missing, remove the
broken <hash> and retry the rename once.
- Treat a
<hash> dir without .complete as absent everywhere (remove before pulling).
- Retry transient rename errors (EIO, EBUSY) a few times with short backoff, and sweep stale
.tmp-* entries at daemon start.
Workaround (what our machine-provider plugin does): before enrolling or resuming a machine,
delete empty runtime/skill-store/<hash> dirs and .tmp-* leftovers; repeat for ~15 min after
enrollment (only .tmp-* older than 10 min).
Environment. bb-app 0.45.0 server and host daemon (daemon installed by the server on a remote
Linux machine, Ubuntu 24.04), claude-code and pi providers. Related but different: #4589 (HTTP 500
on skill-tree fetch, closed).
Symptom.
thread.startfails within seconds with one of:ENOENT: no such file or directory, open '…/runtime/skill-store/<hash>/.last-used'EIO: i/o error, rename '…/runtime/skill-store/.tmp-<hash>-…' -> '…/runtime/skill-store/<hash>'The daemon logs
Failed to pull required injected skill tree … sourceType: data-dir(nameacp-provider) on every start, and it doesn't recover by itself. In the ENOENT caseruntime/skill-store/<hash>/existed but was empty (healthy entries have.complete,.last-used,content/). Deleting that empty dir (and any.tmp-*leftovers) fixed it, and thenext start pulled the tree normally.
Trigger. On our machines the disk is restored lazily after start/resume, so file operations,
including renames, can fail transiently with EIO for a few minutes. The trigger is environmental,
but getting stuck permanently is in the daemon.
What the 0.45.0 daemon does (
daemon-bundle.mjs, the skill-tree install path):<hash>/.completeand<hash>/content; on ENOENT it pulls again. ✔.tmp-<hash>-<pid>-<time>-<rand>/(content,.last-used,.complete), thenrename(tmp, <hash>):<hash>/.last-usedand returning<hash>/contentwithout re-checking<hash>/.complete,so an existing broken
<hash>is never replaced;retry until the next start.
Exactly how the empty
<hash>dir appeared is inferred, not proven (a partial rename, or ahalf-restored dir, then trusted by the EEXIST branch).
Suggested fixes.
<hash>/.complete+content/; if missing, remove thebroken
<hash>and retry the rename once.<hash>dir without.completeas absent everywhere (remove before pulling)..tmp-*entries at daemon start.Workaround (what our machine-provider plugin does): before enrolling or resuming a machine,
delete empty
runtime/skill-store/<hash>dirs and.tmp-*leftovers; repeat for ~15 min afterenrollment (only
.tmp-*older than 10 min).