Problem
The bridge's SQLite database (bridge/state.db) has no mechanism to remove or archive completed job rows. Every webhook delivery that passes through the bridge creates a jobs row carrying the full GitHub webhook payload (10–50 KB of JSON per event). Rows transition to done or failed status but are never deleted.
The claim_keys table similarly grows one row per unique PR and is never pruned.
By contrast, the HUD daemon (bin/hud-daemon.js) already implements a robust retention model: pruneDatabase() runs on a periodic setInterval, deleting events and loaded-file rows older than PROJECT_TTL_MS (168 hours / 7 days by default), with both in-memory ring-buffer caps and SQLite DELETE cleanup. The bridge has no equivalent.
Impact
[PROD] Unbounded disk growth. On a persistent EC2 proxy host (ADR-004), the bridge worker runs continuously. At moderate review volume (20 PRs/week × 2 iterations each), state.db grows ~2 MB/month with no ceiling. Over months of operation, this becomes a reliability risk — SQLite performance degrades as the WAL grows, and the host's disk fills without warning.
[PROD] payload_json retains full webhook payloads indefinitely. Each row stores the complete pull_request_review event including PR body, review comments, and diff metadata. This data has operational value during active processing but none after the job completes. Retaining it indefinitely increases the blast radius of a state.db leak and conflicts with the minimal-data-retention principle.
- No
VACUUM. SQLite does not reclaim disk space from deleted rows without an explicit VACUUM (or auto_vacuum pragma). Even if rows were deleted manually today, the file would not shrink.
Current state
| Component |
Row cleanup |
In-memory cleanup |
Periodic task |
HUD daemon (hud-daemon.js) |
✅ pruneDatabase() with TTL |
✅ pruneInMemoryState() |
✅ setInterval every hour |
Bridge DB (bridge/db.py) |
❌ None |
N/A (no in-memory cache) |
❌ None |
Proposed solution
Add a prune_completed_jobs() function to bridge/db.py and call it from the worker's poll loop at a reasonable cadence (e.g., once per 100 poll cycles or once per hour, whichever comes first).
Retention rules
jobs table: Delete rows where status IN ('done', 'failed') and finished_at < datetime('now', '-7 days'). The 7-day window matches the HUD daemon's PROJECT_TTL_MS default, giving operators time to inspect failures.
claim_keys table: Delete rows that have no remaining jobs rows (all associated jobs have been pruned). This prevents the claim-key counter table from growing indefinitely.
auto_vacuum pragma: Set PRAGMA auto_vacuum = INCREMENTAL on DB init so SQLite reclaims pages progressively. Run PRAGMA incremental_vacuum after each prune pass. This avoids the full-table lock of VACUUM while still reclaiming disk space.
Worker integration
In bridge/worker.py's run() poll loop, call prune_completed_jobs() on a cadence. A simple approach: track an iteration counter and prune every N iterations (e.g., every 1800 poll cycles at 2s intervals ≈ once per hour).
Configuration
BRIDGE_JOB_RETENTION_DAYS (default: 7) — environment variable controlling the retention window, matching the HUD's HUD_PROJECT_TTL_HOURS pattern.
Constraints
Success criteria
Out of scope
- Archiving completed jobs to a separate table or export file (if needed, it's a follow-up).
- Changing the
payload_json column to store a truncated payload (separate concern, could be addressed independently).
- HUD daemon changes (its retention is already adequate).
Problem
The bridge's SQLite database (
bridge/state.db) has no mechanism to remove or archive completed job rows. Every webhook delivery that passes through the bridge creates ajobsrow carrying the full GitHub webhook payload (10–50 KB of JSON per event). Rows transition todoneorfailedstatus but are never deleted.The
claim_keystable similarly grows one row per unique PR and is never pruned.By contrast, the HUD daemon (
bin/hud-daemon.js) already implements a robust retention model:pruneDatabase()runs on a periodicsetInterval, deleting events and loaded-file rows older thanPROJECT_TTL_MS(168 hours / 7 days by default), with both in-memory ring-buffer caps and SQLiteDELETEcleanup. The bridge has no equivalent.Impact
[PROD]Unbounded disk growth. On a persistent EC2 proxy host (ADR-004), the bridge worker runs continuously. At moderate review volume (20 PRs/week × 2 iterations each),state.dbgrows ~2 MB/month with no ceiling. Over months of operation, this becomes a reliability risk — SQLite performance degrades as the WAL grows, and the host's disk fills without warning.[PROD]payload_jsonretains full webhook payloads indefinitely. Each row stores the completepull_request_reviewevent including PR body, review comments, and diff metadata. This data has operational value during active processing but none after the job completes. Retaining it indefinitely increases the blast radius of astate.dbleak and conflicts with the minimal-data-retention principle.VACUUM. SQLite does not reclaim disk space from deleted rows without an explicitVACUUM(orauto_vacuumpragma). Even if rows were deleted manually today, the file would not shrink.Current state
hud-daemon.js)pruneDatabase()with TTLpruneInMemoryState()setIntervalevery hourbridge/db.py)Proposed solution
Add a
prune_completed_jobs()function tobridge/db.pyand call it from the worker's poll loop at a reasonable cadence (e.g., once per 100 poll cycles or once per hour, whichever comes first).Retention rules
jobstable: Delete rows wherestatus IN ('done', 'failed')andfinished_at < datetime('now', '-7 days'). The 7-day window matches the HUD daemon'sPROJECT_TTL_MSdefault, giving operators time to inspect failures.claim_keystable: Delete rows that have no remainingjobsrows (all associated jobs have been pruned). This prevents the claim-key counter table from growing indefinitely.auto_vacuumpragma: SetPRAGMA auto_vacuum = INCREMENTALon DB init so SQLite reclaims pages progressively. RunPRAGMA incremental_vacuumafter each prune pass. This avoids the full-table lock ofVACUUMwhile still reclaiming disk space.Worker integration
In
bridge/worker.py'srun()poll loop, callprune_completed_jobs()on a cadence. A simple approach: track an iteration counter and prune every N iterations (e.g., every 1800 poll cycles at 2s intervals ≈ once per hour).Configuration
BRIDGE_JOB_RETENTION_DAYS(default: 7) — environment variable controlling the retention window, matching the HUD'sHUD_PROJECT_TTL_HOURSpattern.Constraints
DELETEon indexed columns is fast, andincremental_vacuumis non-blocking.db.py, notworker.py.Success criteria
prune_completed_jobs()indb.pydeletesdone/failedrows older than the retention window and orphanedclaim_keysrows.auto_vacuum = INCREMENTALis set ininit_db().test/python/test_bridge_db.pyverifies that completed jobs older than the window are deleted and younger ones are retained.BRIDGE_JOB_RETENTION_DAYSenvironment variable controls the window.Out of scope
payload_jsoncolumn to store a truncated payload (separate concern, could be addressed independently).