Skip to content

DEPLOY HAZARD: after bp-152 the daemon's incompleteness probe collapses to one element and enqueues a backfill on every startup #39

Description

@ascalva

Found in the orchestrator's merge audit of PR #38 (bp-152). Not a defect in bp-152's code —
the builder identified the reader and correctly sequenced its fix to bp-153 (§6's re-home). This
issue records the operational consequence of the gap between the two, because it is a deploy
decision and deploy is the owner's gate.

The mechanism

ops/lifecycle/launcher.py:385-392 (the startup catch-up probe):

store_versions = {(str(r["source_path"]), str(r["digest"]))
                  for r in code_driver.store.all_rows(provenances={Provenance.CODE})}
...
return len(store_versions) < len(ledger)

After bp-152, every code atom row carries source_path="", digest="", title="" (the D1
shed; core/ingest/code_corpus.py:331, stated at :304). So that set comprehension collapses to
exactly one element — {("", "")} — no matter how complete the store is.

1 < 1542 is True, permanently. The probe reports "incomplete" forever and _catchup()
enqueues a backfill on every daemon start.

Why this is the specific thing the probe was built not to do

The probe's own docstring names this failure mode and says a falsifier forbids it:

§6's shorthand distinct digests < distinct versions would false-positive forever — 1,472
distinct blobs < 1,542 distinct (path,blob) pairs even when complete; the falsifier forbids
that loop
, so the probe is like-to-like — finding-0166.

The shed re-creates finding-0166's forbidden loop by a different route: not a wrong comparison, but
a degenerate left-hand side.

Blast radius

Each start enqueues one full backfill (not a tight loop), but a backfill embeds. So the cost is
a full re-embed pass per daemon start, indefinitely, until bp-153's re-home lands. Given the lane's
history of wedging under exactly this kind of unbudgeted work (issue #18, and the Jul 24 battery
drain that wedged code_sync), this is worth not discovering in production.

Disposition

Options, owner's call:

  1. Hold the deploy until bp-153 lands. Cleanest; costs nothing but time, since the live store
    is unchanged in the meantime.
  2. Deploy with [ingestion.code].enabled = false, which gates every path that agent uses
    including startup catch-up (the 2026-07-28 one-section-per-agent split). Lets the rest of the
    daemon move while the code lane waits.
  3. Pull the probe re-home forward out of bp-153 into its own small plan, so bp-152 can deploy
    independently.

Recommendation: (1) unless there is a reason to deploy sooner — bp-153 is the next plan in the
wave anyway, so the window is short.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    route:ownerneeds the ownertrack:code-ingestthe code embed + retrieval pipeline tracktype:defectthe record or the code is wrong

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions