From 32a28e7fb70063ba0a8e238790fefc1f4c944182 Mon Sep 17 00:00:00 2001 From: MXAntian Date: Thu, 1 Oct 2026 15:38:25 +0800 Subject: [PATCH] fix(cli): dedup compact summaries on (source_id, content_hash), not source_id alone MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A compact summary is a running total of the session so far, so the first one is the narrowest. `--store-compact-summary` deduped on source_id alone: the first summary of a session was stored, and every later, wider one hit "already stored" and was dropped. A SessionStart hook normally discards the CLI's stdout, so nothing surfaced the loss. Scale, measured over one hook-trace archive: 190 recorded compactions across 43 sessions. 30 sessions (69.8%) compacted more than once, so 147 summaries (77.4%) could never reach the table however clean the write path was. One session compacted 55 times; 54 of those windows were never in memory at all. Dedup now keys on (source_id, content_hash). Identical content still short-circuits, so a resume replaying the same summary stays a no-op. content_hash and its index already exist (migration 004), and the hash is computed the same way storeMemory computes it for `content`. Rows written before migration 004 have no hash, so an identical replay of one of those is stored once more -- the same thing storeMemory's own dedup does for legacy rows. Checked end to end through the CLI on a scratch DB, one session id: first summary -> stored same content again -> already stored (dedup still works) wider summary -> stored (was "already stored") wider summary again -> already stored Co-Authored-By: Claude Sonnet 5.5 Co-Authored-By: 千夏 --- index.mjs | 11 +++++++++-- 1 file changed, 9 insertions(+), 2 deletions(-) diff --git a/index.mjs b/index.mjs index d220440..196912e 100644 --- a/index.mjs +++ b/index.mjs @@ -4466,9 +4466,16 @@ if (_isMain) { process.exit(1) } const db = getDb() + // Dedup on (source_id, content_hash), not source_id alone. A compact summary is a + // running total of the session so far, so the first one is the narrowest: keying on + // source_id kept exactly that one and silently dropped every later, wider summary of + // a session that compacts more than once. Identical content still short-circuits, so + // a resume replaying the same summary stays a no-op. The hash is the one storeMemory + // writes for `content`. + const contentHash = createHash('sha256').update(String(summary)).digest('hex').slice(0, 16) const existing = db.prepare( - `SELECT rowid FROM memories WHERE source = 'compression' AND source_id = ? AND deleted_at IS NULL LIMIT 1` - ).get(sessionId) + `SELECT rowid FROM memories WHERE source = 'compression' AND source_id = ? AND content_hash = ? AND deleted_at IS NULL LIMIT 1` + ).get(sessionId, contentHash) if (existing) { process.stdout.write(`already stored (rowid ${existing.rowid})\n`) return