diff --git a/.dev-loop/INGEST_REPORT.md b/.dev-loop/INGEST_REPORT.md index 55ccfd1..4f108ac 100644 --- a/.dev-loop/INGEST_REPORT.md +++ b/.dev-loop/INGEST_REPORT.md @@ -1,53 +1,51 @@ -# Knowledge consolidation — 15 open PRs (#17–#40) → one reconciled state - -The 15 open `knowledge/*` PRs (created 2026-08-04 → 2026-08-05, before the -harvest processed-store dedupe fix in #41) contained 123 file-versions of ~75 -unique pages, with the same insight landing at up to 3 different paths across -up to 8 PRs. Per-PR review would re-import those duplicates, so — as with the -#6–#13 consolidation — this branch carries the reconciled end-state and the 15 -PRs are closed in its favor. - -## Verified best-practice - -Every adopted page's sources were carried from its originating PR's flush, where -they were live-verified at flush time; no new URLs were introduced during -consolidation (checked mechanically: every `http(s)` URL in every merged page -appears in a source PR's diff; every added body line in amended pages traces to -a source PR hunk — orphan-line verification). Confidence fields were kept as the -originating flushes set them, except client-side-rate-limiting where the union -of provider-doc citations (Okta, Auth0, GitHub, OpenAI, RFC 6585) supports -`verified` for the load-bearing claims. One subagent's fabricated content (12 -files matching neither main nor any PR, with invented source URLs) was detected -by the same verification and replaced with true PR content. +# Knowledge consolidation — 30 open PRs (#47–#110) → one reconciled state + +The 30 open `knowledge/*` PRs (created 2026-08-06 → 2026-08-17, all after the +harvest processed-store dedupe fix in #41) each carried flush-time dedup against +the then-open PR set, but every branch also rewrote `log.md` and this report and +several branches amended the same pages, so per-PR merge meant a conflict +cascade at every step. As with the #17–#40 consolidation, this branch carries +the reconciled end-state and the 30 PRs are closed in its favor. + +## Method + +Branches were merged chronologically (PR order #47 → #110) onto post-#111 main. +Non-bookkeeping conflicts were union-resolved per file: `related:` lists as id +unions, `last_verified` as the max date, index "load when" rows as the newer +routing framing plus the branch's genuinely new trigger clauses, edge-case +tables as row unions with duplicate-substance rows dropped (one row in +worktree-isolated-workers whose budgeting advice restated an already-merged +row). `log.md` was rebuilt by splicing every branch's own entries into date +order (44 entries added); each branch's added wiki pages were verified present +in the final tree (0 missing). ## Existing-layer check -- Merged-main near-dup scan before consolidation: pairwise Jaccard over - title + "When this applies" across all 141 merged pages → **0 flagged pairs**; - previously merged content carries no duplication. -- Cross-PR dedup during consolidation: 10 duplicate clusters collapsed to one - canonical page each (rate limiting 8→1, call-site enumeration 7→folded into - the canonical merged in #20, stderr/exit-0 diagnostics 4→1, sysroot 2→1, - env-off-switch 2→1, completion predicates 2→1, robots.txt 2→1, - harness-mediated results 2→1, leaked artifacts 2→1, orchestration category - naming unified). Three near-pairs kept distinct after trigger comparison, - with mutual `related:` links (differential setup vs interpretation; expansion - semantics vs off-switch design; import-time tactics vs level choice). -- 24 existing pages received union-merged amendments; additions already present - in main (from #16/#20) were skipped, and all non-canonical `related:` ids - were remapped to canonical page ids (post-merge broken-link scan: 0). +Pages read: testing-quality-tests-that-cannot-fail, testing-quality-behavior-not-implementation, testing-quality-guard-shape-vs-consequence, testing-quality-harness-reverse-controls, testing-quality-policy-at-several-return-sites, testing-data-test-data-and-isolation, backend-common-change-impact-call-site-enumeration, backend-common-change-impact-inserting-a-guard-before-an-existing-side-effect, backend-common-change-impact-cross-module-consumer-census, backend-java-jpa-raw-jdbc-inside-a-jpa-transaction, infrastructure-agent-orchestration-worktree-isolated-workers, infrastructure-agent-orchestration-shared-run-state, infrastructure-agent-orchestration-session-completion-gates, qa-document-verification-spec-document-gates, qa-document-verification-editing-a-gated-document, qa-exploratory-guard-true-path-coverage, qa-environments-headless-browser-bot-blocking, qa-process-defect-class-resweep-after-review, qa-process-scope-purity-checks (conflict-resolution and qualifier-polish reads on the final tree) + +- Flush-time dedup notes were honored as recorded in each PR title (e.g. #74's + two insights folded into #73/#52's pages arrive via those branches; #86's + duplicate dropped); no dropped insight was re-imported. +- Where a branch's index row diverged from a later routing revision already on + main (the 2026-08-12 disjoint-scope revision for spec-document-gates vs + testing/quality), the merged row keeps the revision's scope and carries only + the branch's new trigger clauses. +- Post-merge scans: 0 conflict markers, broken-link scan and index-coverage + lint run on the final tree (see log entry). + +## Open-PR check + +This consolidation *is* the reconciliation of the open-PR backlog: all 30 open +`knowledge/*` PRs (#47–#110) are merged into this branch and closed in its +favor. No other open knowledge PRs remain; #111 was merged to main first and +this branch is based on it. ## Routing decision -- New categories: `infrastructure/agent-orchestration` (5 pages; unified the - competing `orchestration`/`agent-orchestration` names), `databases/data-survey` - (1), `qa/deliverables` (1). All other pages route into existing categories. -- Canonical-path decisions: rate limiting → `backend/common/reliability/` - (sits beside timeouts-and-retries; 6 of 8 variants chose it); stderr - diagnostics → `platforms/processes/` (concern spans beyond shells); leaked - artifacts → `testing/data/artifact-leakage-from-a-suite`; call-site - enumeration → the existing `backend/common/change-impact/` page. -- All 38 new pages listed in their domain indexes (nearest-index rule; backend - routes via its python sub-index for bytecode-cache-staleness); INDEX.md domain - summaries updated for infrastructure/qa/databases. Full-wiki lint: frontmatter, - ids, related-links, index coverage, size, qualifiers, staleness → 0 findings. +- New categories: `backend/common/ml` (mape-aligned-point-prediction), + `backend/python/packaging` (data-files-and-install-paths), + `security/incident-response` (verifying-assumed-security-agents, + process-identity-by-path-and-hash). All other 53 new pages route into + existing categories; 51 existing pages received union-merged amendments. +- All new pages are listed in their domain indexes (nearest-index rule); + INDEX.md domain summaries updated for backend, frontend, security. diff --git a/INDEX.md b/INDEX.md index 838086d..8f59d1a 100644 --- a/INDEX.md +++ b/INDEX.md @@ -12,13 +12,13 @@ follow the cross-pointers in their index or take the next matching seeded domain | Domain | Status | Route here when | |--------|--------|-----------------| | [databases](wiki/databases/index.md) | **seeded** | Designing schemas/tables/keys, choosing or evaluating indexes, writing or optimizing queries, choosing transaction/isolation behavior, surveying live data to derive a rule, verifying additive migrations | -| [backend](wiki/backend/index.md) | **seeded** | Server-side application code — language-agnostic (`common/`: API contracts, call-site enumeration before a contract change, idempotency, JWT, timeouts/retries, caching, jobs, transactions in app code, shared state/pools, errors, consuming LLM APIs (completion validation, context budgeting), consuming external-API responses, externally-owned defaults, object-storage references, sync-vs-async integration choice, WebSocket/SSE connection lifecycle) plus stack subtrees: `java/` (JPA, Spring proxies, JVM threads/memory), `node/` (event loop, promises, runtime validation, shutdown), `python/` (GIL/asyncio, pydantic, WSGI/ASGI workers, language traps) | +| [backend](wiki/backend/index.md) | **seeded** | Server-side application code — language-agnostic (`common/`: API contracts, call-site enumeration before a contract change, idempotency, JWT, timeouts/retries, caching, jobs, transactions in app code, shared state/pools, errors, consuming LLM APIs (completion validation, context budgeting), MAPE-aligned point-prediction calibration, consuming external-API responses, externally-owned defaults, object-storage references, sync-vs-async integration choice, WebSocket/SSE connection lifecycle) plus stack subtrees: `java/` (JPA, Spring proxies, JVM threads/memory), `node/` (event loop, promises, runtime validation, shutdown), `python/` (GIL/asyncio, pydantic, WSGI/ASGI workers, language traps, packaging data files with `importlib.resources`) | | [frontend](wiki/frontend/index.md) | **seeded** | Web UI code: state placement, rendering performance, in-UI data fetching (races, infinite scroll), auth token handling, forms, XSS-safe output, accessibility, agent-facing tool surfaces (WebMCP) | | [infrastructure](wiki/infrastructure/index.md) | **seeded** | CI/CD pipelines, secrets in build/deploy, container image builds, rollout/rollback strategy, observability (logs/metrics/alerting), per-environment/path-valued config, multi-agent orchestration (worker liveness signals, shared run state, tmux pane delivery, completion gates, worktree-isolated workers) | | [testing](wiki/testing/index.md) | **seeded** | Writing or structuring automated tests: level choice, cases/assertions, test data, mock decisions, flaky tests (release-process quality → qa) | -| [qa](wiki/qa/index.md) | **seeded** | Release-quality process: release gates, regression scoping, bug reports, severity/priority triage, exploratory testing (guarded-path coverage, override matrices), scope-purity gates, sourcing deliverable documents from generated artifacts, automated verification of document deliverables (spec/RFC gates) (writing automated test code → testing) | +| [qa](wiki/qa/index.md) | **seeded** | Release-quality process: release gates, regression scoping, bug reports, severity/priority triage, exploratory testing (guarded-path coverage, override matrices), scope-purity gates, sourcing deliverable documents from generated artifacts, verifying the quantitative claims in a document before publishing it, automated verification of document deliverables (spec/RFC gates) (writing automated test code → testing) | | [debugging](wiki/debugging/index.md) | **seeded** | Diagnosing a failure — finding what is wrong and why: reproducing, bisection, hypothesis testing, traces/logs, intermittent failures (fixing the diagnosed fault → its owning domain) | -| [security](wiki/security/index.md) | **seeded** | Trust-boundary decisions: input validation, session-vs-token auth choice, per-resource authorization (IDOR), secrets hygiene, dependency trust, PII handling, in-session agent tool exposure (prompt-injection blast radius) (XSS rendering → frontend; CI secrets → infrastructure; JWT implementation → backend/frontend auth) | +| [security](wiki/security/index.md) | **seeded** | Trust-boundary decisions: input validation, session-vs-token auth choice, per-resource authorization (IDOR), secrets hygiene, dependency trust, PII handling, in-session agent tool exposure (prompt-injection blast radius), the author identity a commit publishes to a public repository, host-compromise triage / incident response (verifying assumed security agents, identifying masquerading processes) (XSS rendering → frontend; CI secrets → infrastructure; JWT implementation → backend/frontend auth) | | [platforms](wiki/platforms/index.md) | **seeded** | OS-level differences breaking code across macOS/Linux/Windows: shell portability, BSD-vs-GNU CLI, filesystem case/line endings, Unicode normalization in text/file-name matching, commands inspected before execution, background services/cron, invoking prompt-capable CLIs non-interactively, toolchain version pinning | | [mobile](wiki/mobile/index.md) | **seeded** | App-side iOS/Android/cross-platform: process death/state survival, offline-first sync, mobile-network calls, store rollout/hotfix strategy, startup time | diff --git a/log.md b/log.md index 96674fa..e2b0e9e 100644 --- a/log.md +++ b/log.md @@ -43,6 +43,51 @@ Append-only. Format: `## [YYYY-MM-DD] = n` / `toHaveLength(n)` count assertion over a call appearing at several sites stays green when the one site the guard was written for is deleted — enumerate the sites and bind each to a bounded order anchor that occurs exactly once in the file, or to a function-body slice; prove each by deleting only its own site, and run a reformat control. Measured: greedy vs lazy quantifiers give identical verdicts, so the bound and the anchor's uniqueness are what constrain the match), frontend/data-fetching/query-state-vs-fetch-state (a `data | undefined` component prop collapses TanStack Query's two orthogonal axes; a disabled or offline-paused query is `status: pending` with `isLoading === false` and `isError === false`, so "undefined means loading" renders a spinner no fetch will resolve — pass status+fetchStatus or an explicit union and test one case per cell), backend/python/language/default-encoding-in-text-io (a byte round-trip cannot prove an `encoding=` fix on a UTF-8 locale — run the real entry point under `-X warn_default_encoding -W always::EncodingWarning` and assert zero warning lines naming that file). Merged into existing: tests-that-cannot-fail (whole-suite edge row now routes surviving mutants through classification instead of reading them all as missing tests), harness-reverse-controls, behavior-not-implementation, guard-shape-vs-consequence, async-ui-states (+disabled/paused edge row), bytecode-cache-staleness, timezone-and-locale — related links both ways. All cited URLs opened this session; two local reproductions (CPython 3.14.6 EncodingWarning discriminator vs byte-identical round-trip; `@tanstack/query-core@5.100.14` queryObserver.js:308-332 `isLoading = isPending && isFetching`). +## [2026-08-07] revise | testing/quality/guard-shape-vs-consequence — corrected a misattributed citation found by the pre-PR adversarial pass: the sentence "you cannot safely refactor code if you know you need to adapt the tests afterwards to get them passing again" was presented as the Google Testing Blog article's own, but re-fetching the page shows it is a reader comment (2015-02-04) with different wording ("refactor stuff", "know for sure"). The article body was not retrievable in full, so the bullet now cites the URL for the change-detector category without quoting it, and states the correction inline. The same quote had been copied into a new page in this flush before verification — the lesson being that a citation already present in the wiki is not a verified citation. +## [2026-08-07] ingest | backend-common-ml-mape-aligned-point-prediction — NEW category backend/common/ml (no existing category covers training/evaluating predictive models; llm covers only consuming LLM APIs). A median-predicting regression model (log target + L1) scored by MAPE is a structural overpredictor: the MAPE Bayes rule is the 1/y-reweighted median (Gneiting arXiv:0912.0902 Table 5; de Myttenaere arXiv:1605.02541), = median × exp(−σ²) under lognormality. Directive: per-row `pred × exp(−λσ²)` with σ from q16/q84 quantile spread, λ selected from 0.5 by all-periods holdout improvement (theory λ=1.0 overcorrects because spread-estimated σ is inflated). Theory verified vs sources; λ practice field-tested (Seoul AVM, 4/4 holdout years improved at λ=0.5, 2/4 degraded at λ=1.0). Two queued dev-loop orchestration candidates dropped as pending-duplicates of open PR #51. +## [2026-08-07] ingest | knowledge-flush of 3 queued insights: 1 merged, 2 dropped as in-flight duplicates. Merged: platforms/processes/tool-diagnostics-without-a-failing-exit-code — the warnings-as-errors adoption check gains a fourth control input (a valid file that legitimately warns: deprecations, accept-and-warn declarations); without per-diagnostic severity control the promotion switch and the intentional-warning feature are mutually exclusive, recorded as a platform defect (+1 Do-5 caution, Do-6 fourth state, +1 edge case, +1 Instead-of; sources: GCC -Werror=/-Wno-error= granularity, rust-unofficial deny-warnings anti-pattern, lnpl 0.2.0 --strict field reproduction rc=2 on a legitimate `on schedule` declaration). Cross-linked both ways with backend/common/api-design/unenforced-declarations (its accept-and-warn shape is exactly the diagnostic that collides with a blanket gate). Dropped: worktree_escape read-only escalation round-trip (same evidence already in open PRs #47 and #51 on worktree-isolated-workers) and orca dispatch-binding stage taxonomy (already in open PR #51 on pane-delivery-confirmation, incl. the same three field observations). +## [2026-08-07] ingest | knowledge-flush: 1 new page testing/strategy/signal-delivery-to-a-process-under-test — POSIX §2.11 sets SIG_IGN for SIGINT/SIGQUIT on `&` jobs when job control is off, so `kill -INT` in a script harness is discarded and `wait` hangs; deliver via a subprocess driver (Popen.send_signal), `set -m`, or SIGTERM. Verified against POSIX + bash §3.7.6 + local sh/bash/zsh reproduction. Two sibling candidates (worktree_escape read-only escalation budgeting; orca bind-stage branching) dropped as pending duplicates of open PR #51. +## [2026-08-07] ingest | knowledge-flush of 3 queued insights → 1 ingested, 2 dropped as in-flight duplicates of open PR #51. New: backend/common/change-impact/corpus-sweep-before-a-rejection-rule (a rule that will start rejecting silently-accepted input is bounded by a throwaway accept/reject predicate swept over the whole corpus *before* production code — the plan carries the reject count, the rejected-path list, and the enumeration method, and the list must equal the known-defective set; re-run the same script after landing and require the verdict sets to agree). Sourced against three ecosystem implementations of the same method (eslint-remote-tester's "not enough to test the rule only against unit tests and a small amount of repositories", clippy lintcheck's recorded-and-diffed warning logs, rust crater's two-toolchain corpus comparison); the pre-implementation ordering is the field-tested refinement (linkly #53: 148 sources swept → 2 rejects, one false positive caught at plan time, `.replace()`-assembled fixtures a stated blind spot). Reciprocal related-links added to change-impact/call-site-enumeration and qa/process/regression-scope. Dropped: the guardrail read-only-escalation insight and the pane-binding failure taxonomy — PR #51 already carries both in more precise form (it corrects the first: a pure read does not fire; the rule needs a write verb or an absolute redirect in the same command). +## [2026-08-07] ingest | knowledge-flush of 4 queued insights → 2 ingested, 2 dropped as in-flight duplicates of PR #51. New: testing/data/harness-vs-run-path-fixtures — a spec harness and the production entry point each synthesize the input and the harness reads a narrower declaration set, so a guard whose operand is absent lands on the not-true branch and the run reports "step skipped, exit 0" as if it were program behavior; give both entry points one synthesizer, make any narrowing an explicit author-overridable argument, and separate absent-operand from false-operand before filing. Merged: testing/quality/spec-artifact-checks +Do step 3 / +3 edge cases / +1 Instead-of — a negative control that quotes the whole target row (prose cells included) stops matching after an ordinary rewording, `str.replace` returns the unchanged document rather than raising, the check then passes correctly against the original, and the control reports green while mutating nothing; anchor on the row's key-column prefix and assert `mutant != original` (or require `re.subn`'s count). Sources live-verified: PostgreSQL functions-comparison (null comparison yields unknown, `NULL::boolean IS TRUE` → f), jq manual (absent key → null; if takes else for false or null), Python stdtypes `str.replace` / re `subn`, ISTQB branch-coverage. Two local reproductions (Python 3.14.6 + jq): absent operand and literal false select the same branch; a stale literal anchor leaves `mutant == doc` with no exception while the prefix anchor applies (`re.subn` 0 vs 1). +## [2026-08-07] dedup | Dropped 2 pending candidates as duplicates of open PR #51 (knowledge/dch0202-20260806-183029), which carries both in better-sourced form: the worktree guardrail read-only escalation candidate (PR #51's worktree-isolated-workers rows already cover budgeting the escalation round trip and stating "reads are approved" in the worker's first brief — and its local reproduction corrects the candidate's premise: plain reads of a sibling worktree pass, and the rule fires only when a write verb or an absolute-path redirect co-occurs in the same command string); and the orca terminal-binding candidate (PR #51's pane-delivery-confirmation rows already cover idle-prompt check before binding, worker-done as a report rather than a turn end, runtime-unavailable → wait and bind a fresh unit, agent-unconfigured → close the pane and create a worker-mode agent, and passing the worktree alongside the pane). Same pair previously dropped by flushes #56, #57 and #58. +## [2026-08-07] ingest | knowledge-flush of 3 queued candidates → 1 ingested, 2 dropped. New: platforms/toolchains/environment-resync-removes-undeclared-packages — a dependency command that resyncs the environment deletes packages absent from the lockfile, so an ad-hoc `uv pip install` dev tool disappears on an unrelated dependency change; declare dev tools in a PEP 735 dependency group (synced by default), and gate only the *exact* commands on background jobs finishing. The harvested candidate asserted `uv add`/`uv remove` were symmetric; local reproduction (uv 0.11.5, macOS) refuted that — `uv add` left the undeclared package installed, `uv remove` deleted it twice, `uv remove --no-sync` preserved it, and `--help` shows only `uv sync` exposes `--inexact` and only `uv run` exposes `--exact`, so `uv remove` has no exactness escape hatch. Also reproduced the delayed-failure mechanism: a process that had already imported the package kept running off the module cache while a later new import in the same process raised ModuleNotFoundError. Sources live-verified (docs.astral.sh sync/dependencies/CLI reference, PEP 735). Plumbing: platforms/index.md toolchains row + reciprocal related-links on version-management and background-services. Dropped as in-flight duplicates of open PR #51: the guardrail worktree_escape read-vs-write candidate (whose naive "reads escalate" claim #51 already corrects with a local reproduction) and the Orca dispatch-binding taxonomy candidate (covered by pane-delivery-confirmation and worktree-isolated-workers). Both are recurring re-emissions — already present in the processed queue 15x and 10x respectively. +## [2026-08-08] ingest | knowledge-flush of 3 queued insights → 1 ingested, 2 dropped as in-flight duplicates. New: infrastructure/agent-orchestration/unattended-worker-questions (a worker agent's in-band question reaches nobody — with a TTY the chooser waits indefinitely, without one it self-answers empty in ~37 ms; classify a live-terminal stall from the terminal tail before restarting, unblock with an allowlisted key sequence validated whole and re-driven until the idle prompt returns, re-send the prompt the chooser swallowed, and prevent recurrence with a durable out-of-band question record the watcher wakes on). Sources: anthropics/claude-code#50728 (closed as not planned) and #29530 (open), tmux(1), plus the shipped dev-loop 1.4.2 orchestrate scripts (ask-coordinator.sh, watch-status.sh, send-prompt.sh keys, orca-worker-stalled.sh's 75-minute measurement) and a field observation on worker lo-4-qag1. Dropped: the worktree_escape read-only-escalation candidate (already in flight twice — PR #47 and PR #51) and the Orca dispatch-binding taxonomy candidate (already in flight — PR #51's pane-delivery-confirmation rows and PR #47's control-signals row). +## [2026-08-08] contradiction | PRs #47 and #51 give opposite mechanisms for the same guardrail. #47's worktree-isolated-workers row claims "a conservative rule treats any cross-worktree path reference in command text as a potential write, reads included" and that the read/write asymmetry is version-dependent; #51's row claims reads still pass alone and fire only when a write verb or an absolute redirect appears anywhere in the same command string. Read against groundwork guardrails `plugins/guardrails/hooks/bash-guard.sh` (worktree_escape), #51 is correct: after stripping the worker's own worktree path and any configured allowPaths, a surviving `$main_root/` mention fires only if `match "(^|[[:space:];&|(])(rm|mv|cp|tee|mkdir|touch|install|dd)[[:space:]]"` or `match "(>|>>)[[:space:]]*[\"']?/"` also matches — the two tests are independent, which is why a read sharing a command line with any write verb escalates. Resolve in favour of #51's wording when merging. +## [2026-08-08] ingest | knowledge-flush of 5 queued candidates → 2 promoted, 3 dropped as in-flight duplicates. New: qa/deliverables/command-transcripts-in-a-document (re-capture a CLI transcript at paste time and give each stream its own labelled block — a `2>&1` capture records the environment's flush order, not the program's write order; local three-way reproduction on CPython 3.13/macOS: through a pipe the merged order was all-stderr-then-all-stdout, under `python -u` and on a pty it was the program's order, one unchanged binary; POSIX stdin.html — stderr "shall not be fully buffered", stdout fully buffered iff not on an interactive device; doctest captures stdout but not stderr). Merged: qa/document-verification/spec-document-gates +5th "External agreement" axis / +4 edge cases / +1 Instead-of (when a documented table's cells are copies of a code constant, resolve the owning symbol and compare per row — asserting the invariant the table itself states makes the document its own oracle, so the gate built to catch drift is what pins the stale claim green; reproduced 2026-08-08 in `linkly`: `docs/ENFORCEMENT-MATRIX.md` §C reads `warning` in all five rows while `SEVERITY_OF` grades three `info`). Reciprocal related-links added on generated-artifacts-as-deliverable-source, tool-diagnostics-without-a-failing-exit-code, non-interactive-cli-invocation. +## [2026-08-08] dedup | Dropped 3 orchestration candidates as pending duplicates of open PRs, not re-ingested: worktree_escape fires `ask` on read-only cross-worktree access and the coordinator must budget the escalation round trip → already in #51's worktree-isolated-workers hunk (identical directive, richer reproduction); the Orca dispatch-binding taxonomy (runtime_unavailable = occupied turn tail → wait and rebind; agent_unconfigured = dead agent → replace; always pass the worktree with the pane) → already in #51's pane-delivery-confirmation hunk; a tmux worker wedged on a numbered in-band chooser → already in #64's unattended-worker-questions (classify by terminal tail, unblock by allowlisted key, re-send the in-flight prompt). Merging #51 and #64 retires this recurring trio. +## [2026-08-08] lint | 0 errors fixed, 0 reported — changed-page pass over qa/deliverables/command-transcripts-in-a-document (new, 62 body lines, verified with 4 cited sources), qa/document-verification/spec-document-gates (85 body lines), and the three pages that gained reciprocal related-links. Checks run: sources-vs-confidence, prohibition-outside-Instead-of, related/inline id resolution, domain-index presence and trigger agreement, vague qualifiers, body length. +## [2026-08-09] ingest | knowledge-flush of 5 queued insights → 2 new pages, 3 dropped as in-flight duplicates. New: backend/common/errors/diagnostics-from-a-shared-code-path (a rejection message emitted from a path two constructs share takes its subject as a caller-supplied parameter, and its *repair* is checked by executing it from each emitting path — the subject fails visibly, the repair does not; rated like rustc `Applicability`, asserted once per emitting path). New: qa/deliverables/exclusivity-and-absence-claims (falsify an "only way" / "cannot be expressed" / "exactly N forms" claim before writing it, and publish the rule that generates the forms rather than the enumeration — an enumeration records what the author knew and goes stale when a producer is added). Sources live-verified this session: rustc-dev-guide diagnostics + `rustc_errors::Applicability`, NN/g error-message guidelines, Dijkstra EWD303, SEP Popper, Google docguide. Dropped (already carried by open PRs, equal or better form): guardrail read-only worktree escalation → #47/#51 (worktree-isolated-workers), Orca terminal/dispatch binding taxonomy → #51 (pane-delivery-confirmation), tmux in-band question menu → #64 (unattended-worker-questions). Reciprocal related-links added to 6 existing pages. +## [2026-08-09] ingest | knowledge-flush of 3 queued insights → 1 ingested, 2 dropped as in-flight duplicates of open PR #47. New: qa/document-verification/generated-reference-drift-gates — a closed-vocabulary reference an agent reads back (DSL verbs, config keys, diagnostic codes) is generated from the owning constant, gated by a suite-called `--check` mode, and carries a machine-readable generated banner; the page's distinct claim is that `--check` alone is insufficient because it compares the committed file to the generator, so a generator that hardcodes a column regenerates green — per-member coverage assertions against the constant plus a not-a-single-repeated-value negative control are a separate layer. Sources live-verified: go generate's `^// Code generated .* DO NOT EDIT\.$` convention (pkg.go.dev/cmd/go), prettier `--check` exit 1 for CI, Google docguide "same CL"/"link to it instead", Anthropic define-tools "by far the most important factor". Local reproduction (linkly lnpl 0.3.0): `gen_plugin_references.py --check` rc=0 in-suite with hand-edit/missing-file cases asserting rc=1, and the counter-example `examples/login.lnpl` compiling rc=0 with 3 `unknown-verb` no-ops. Back-linked from spec-document-gates and generated-artifacts-as-deliverable-source. Harvested "platforms" domain hint re-routed to qa/document-verification (the case is a document gate, not an OS-level difference). +## [2026-08-09] dedup | Dropped 2 queued candidates as pending duplicates of open PR #47 (head knowledge/dch0202-20260806-130040), which already carries both as edge-case rows on infrastructure/agent-orchestration/control-signals-vs-primary-artifacts with the same 2026-08-06 field evidence: (a) simultaneous worker silence caused by a CLI usage-limit pause that passes every liveness check — find the `You've hit your session limit · resets HH:MM` marker, resume after the reset with a state-recheck → remaining-DoD → completion-signal prompt; (b) a dispatch issued immediately after `worker_done` failing `runtime_unavailable` and consuming the task — wait for the substrate's `tui-idle` signal and create a new task from the same spec. No sibling PR opened. +## [2026-08-10] ingest | Folded one queued insight into this PR's page rather than opening a sibling: testing/quality/source-text-wiring-assertions +step 2 (make the assertion's subject the comment-stripped file, taking comment ranges from the language's tokenizer — ESLint: "While comments are not technically part of the AST", so a rule reaches them via `sourceCode.getAllComments()` instead of node traversal), +4 edge-case rows (comment-only change reddens the guard; negative `not.toMatch` and count assertions are the shapes comments flip; no parser available; the code under test is itself about comments), +3 Instead-of rows, and a step-6 note that with step 2 the comment control is green by construction so a red one indicts the stripper. Verification corrected the queued directive: measured 2026-08-10 in Node, `src.replace(/\/\*[\s\S]*?\*\//g,'').replace(/\/\/.*$/gm,'')` truncates `"https://api.example.com//v2/items"` to `"https:` and a string-aware variant still truncates `/a//b/` to `/a` — so the page prescribes the tokenizer and records the regex forms only as Instead-of rows with that measurement. Field basis: three guards in one rtb-unified PR reddened on correct code from a JSDoc, a `Number.isFinite` comment, and a `SOURCE.slice(...)` comment. Independent adversarial review then corrected the fold itself: `ts.createSourceFile` yields `comments === undefined` and no comment nodes (so it strips nothing — step 2 now names `getLeadingCommentRanges`/`espree`/`@babel/parser`), the step-6 "green by construction" claim covered only the comment half (a reformat still reddens the bound), the no-parser fallback had to cut each line to EOL because filtering whole comment lines misses trailing comments, and N must be measured on the stripped subject. +## [2026-08-10] ingest | knowledge-flush of 3 queued insights → 2 new pages, 1 folded into open PR #52. New: backend/common/integrations/consumer-required-fields (an adapter built from the consumer's docstring cannot tell required from optional — call the real consumer once with a mapped record, split the fields it reads into loud (presence check / subscript → raises on record 1) and silent (read with a default → wrong value, no error), and give each silent field a two-run assertion that drops it and requires the output to differ; sources: docs.pact.io "check that all the calls to your test doubles return the same results as a call to the real application would" + "unlike a schema or specification (eg. OAS), which is a static artefact"; json-schema.org "By default, the properties defined by the `properties` keyword are not required"; Python reproduction 2026-08-10 — missing `assignee_id` raised ValueError, missing `desc` moved the score 21.0 → 1.0 silently). New: infrastructure/observability/suppression-state-and-delivery-failure (write the cooldown/sent mark only after a send that reported success, give the send its own exit status, make it injectable, and assert the succeeding-sender and failing-sender worlds as two tests; sources: Alertmanager `RetryStage` "notifies via passed integration with exponential backoff until it succeeds" then `SetNotifiesStage` "sets the notification information about passed alerts. The passed alerts should have already been sent to the receivers." with `DedupStage` filtering "based on a notification log"; prometheus-operator Watchdog runbook "an alert meant to ensure that the entire alerting pipeline is functional" for the external heartbeat; field measurement rtb-mac-server-k8s bin/gitops-deploy.sh where the always-failing stub had fixed the pre-send ordering as the expected contract — the correct fix turned 3 tests red). Folded (not re-ingested): the comment-vs-code source-assertion insight went to open PR #52's testing/quality/source-text-wiring-assertions rather than a sibling page; verification there found the queued directive's own strip regex defective — measured 2026-08-10 in Node, `s.replace(/\/\*[\s\S]*?\*\//g,'').replace(/\/\/.*$/gm,'')` truncates `"https://api.example.com//v2/items"` to `"https:` and a string-aware variant still truncates a regex literal containing `//`, so the page takes comment removal from the language's tokenizer with a control, not from a regex. +## [2026-08-10] ingest | knowledge-flush of 3 queued insights, 3 new pages (no merge target had the same trigger). testing/mocking/captured-call-arguments — a spy that stores only the argument a review flagged leaves the same call's siblings unasserted; record the call whole (assert_called_with / toHaveBeenCalledWith / all-args matchers), bind the stub to the real signature with spec so positional-vs-keyword is one claim, split "the constant is right" from "the call site passes it on", and re-prove per argument after extracting a resolver (which moves the gap up a layer rather than closing it). Reproduced locally 2026-08-10: partial capture stayed green on a host 0.0.0.0→127.0.0.1 mutation while the whole-call assertion reddened and stayed green on the unmutated call. databases/indexing/trigram-index-short-patterns — a pg_trgm wildcard segment under 3 characters yields no extractable trigrams, so GIN degenerates to a full-index scan while the plan still reads Bitmap Index Scan; branch on whether short keywords are supported input (minimum-length policy / a selective driver index / pg_bigm for LIKE-only workloads / case-folded expression index when ILIKE is required). backend/java/jpa/raw-jdbc-inside-a-jpa-transaction — @Transactional(timeout=N) reaches a raw JdbcTemplate only through a ConnectionHolder bound by JpaTransactionManager (needs the DataSource *and* a JpaDialect that exposes the JDBC connection, same DataSource instance), and the two paths raise different Spring exceptions: Hibernate→QueryTimeoutException vs raw JDBC→DataAccessResourceFailureException on Spring Framework ≤6.2.x. Version boundary verified across tags: the "57014" special case is absent in v5.3.31/v6.0.0/v6.2.0/v6.2.1/v6.2.3/v6.2.5/v6.2.8 and present in v7.0.0. Plumbing: testing, databases and backend/java indexes each +1 row; reciprocal related links added on what-to-mock (+ outbound-contract row pointer), tests-that-cannot-fail, index-selection (+ text-search row pointer), timeouts-and-retries, transaction-boundaries. +## [2026-08-10] ingest | knowledge-flush, 4th candidate of the same flush (it surfaced late: the queue was counted with `wc -l`, which reported 3 because one session file's last line had no trailing newline; the JSON-parsing retirement step moved 4). New: testing/quality/default-values-under-test — a default named as a number in a spec (`ttl_s=600`, `max_tokens=256`) is guarded by neither a mechanism test that passes the value in nor a defaults-constructing test that never exercises it; push the default to its observable point (TTL → both sides of the boundary, cap → exactly the cap then one more) and require red in both mutation directions. Reproduced 2026-08-10 (Python 3, TTL+cap store, 4 pre-existing tests): `ttl 600→1` GREEN and `ttl→6000`/`cap→9999` GREEN against the existing suite, `cap 256→1` caught incidentally, all four RED after adding two bidirectional default tests, unmutated suite green as the control. That sharpened the source candidate, which stated the survival as flat rather than direction-dependent: the grow direction is what no incidental test catches, since a test issuing N items passes for every cap ≥ N and one consuming immediately passes for every TTL > 0. Kept separate from minimum-case-set (case set for the behaviour, not for a shipped default) and from the in-flight #49 unasserted-return-fields (return direction) — reciprocal related links added to minimum-case-set, tests-that-cannot-fail and captured-call-arguments; testing index +1 row. +## [2026-08-11] revise | Folded one queued insight into this PR's page rather than opening a sibling: backend/python/language/bytecode-cache-staleness — the `-B`/`PYTHONDONTWRITEBYTECODE` edge row is corrected (both options govern *writing* only, per docs.python.org/3/using/cmdline.html, so a `.pyc` from an earlier run is still read and the failure survives; purge `__pycache__` once before the run and keep `-B` for the rest), and a new row covers the fresh-spec import path (`spec_from_file_location` + `module_from_spec` + `exec_module` bypasses `sys.modules`, not the on-disk cache; `exec(compile(text, path, "exec"), ns)` consults no cache). Measured 2026-08-11 on Python 3.14.6/macOS: a same-length rewrite under a re-pinned mtime returned the stale value through the fresh-spec loader, and again under `python3 -B` with the cache present. +## [2026-08-11] ingest | knowledge-flush of 3 queued insights → 1 new page, 2 folded into open PRs. New: databases/query-optimization/comparing-two-execution-plans — a two-arm `EXPLAIN (ANALYZE)` comparison whose fast arm shows `never executed` is confounded (X moved *and* whether rows reached the expensive subtree moved), so attribution needs a third arm holding X at the fast setting with rows flowing; `never executed` is emitted whenever a node has no instrumentation with a positive `nloops` (the `else if (es->analyze)` branch also fires when `planstate->instrument` is NULL, so it is not a biconditional on `nloops == 0`), and only in TEXT format — JSON/XML/YAML emit `Actual Rows`/`Actual Loops` of 0 unconditionally and the time fields only under `es->timing`, so a parser keyed on the time scores a skipped subtree as 0 ms. Corrected the harvested candidate's claim that a client-cancelled arm's duration is recoverable from `pg_stat_statements`: the view updates "only for successful operations", so a cancelled execution contributes nothing — recover by re-running without the client deadline, or by reading `pg_stat_activity.query_start` while the statement is still running. Sources live-verified and pinned to commit `083ac03` (postgres explain.c — `REL_17_STABLE` is a branch, not a tag): the nloops branch was read directly, plus the pg_stat_statements, monitoring-stats, using-explain and ddl-partitioning docs (the last is the only official page that names `(never executed)`). An independent adversarial cross-check corrected three claims before merge: `auto_explain` cannot capture a cancelled statement either (same `ExecutorEnd_hook` blind spot), the `es->timing` guard on the emitted time fields, and the `pg_stat_statements` column name following the installed extension version rather than the server version. Cross-links added both ways with reading-execution-plans. +## [2026-08-11] dedup | 2 of 3 queued candidates folded into in-flight PRs rather than re-ingested (open-PR check, skill step 2b′). (a) Spring per-technology query-timeout attribution → PR #73 backend/java/jpa/raw-jdbc-inside-a-jpa-transaction: same production incident and the same measured durations (10,012 / 151,558 / 163,489 ms) already in that page; unique additions pushed to that branch — MyBatis as a third access technology whose `defaultStatementTimeout` is applied by `BaseStatementHandler.setStatementTimeout` independently of any bound transaction, the duration-fingerprinting rule (a duration matching no configured value and varying run to run means no timeout is applied on that path), and the message-vs-SQLSTATE distinction (`canceling statement due to user request` is the external-cancel fall-through branch; statement timeout carries the same SQLSTATE 57014 with different message text). (b) Mutation-harness `__pycache__` staleness → PR #52 backend/python/language/bytecode-cache-staleness already carries the (mtime, size) mechanism, equal-size mutations, cache purge and mtime bump; unique addition pushed to that branch — `shutil.copy2` names itself as an mtime-preserving restore, `copyfile`/`copy` do not. ## [2026-08-12] revise | routing: disjoint scopes for the doc-gate cluster (testing/quality ↔ qa/document-verification) and the flaky pair (testing/flaky ↔ debugging/concurrency); INDEX backend LLM phrasing; databases→backup cross-pointer (#37) +## [2026-08-12] ingest | knowledge-flush fold into this branch (PR #64). New: infrastructure/agent-orchestration/usage-limit-paused-workers — workers billed to one seat exhaust one shared allowance and stop together while every liveness check passes; classify from the marker (`You've hit your session/weekly/Opus limit · resets …`), because session and weekly are shared across all models while only the Opus limit is cleared by `/model`; wait for the stated reset, then resume with a prompt naming the state re-check, the remaining done-criteria, and the completion signal, since a bare "continue" across the boundary re-reads and misinterprets the plan and redoes finished work. Expands the one-row usage-limit case in this branch's unattended-worker-questions classification table rather than duplicating it. Sources: code.claude.com/docs/en/errors (marker strings verbatim, block-until-reset, model-sharing), code.claude.com/docs/en/costs (per-seat rolling five-hour + weekly allowance shared with Claude chat and Cowork; "a single burst of heavy activity, such as a large workflow fanout, can exhaust the weekly allowance"), anthropics/claude-code#5977 (context loss after reset), #36320 (auto-resume still unimplemented), plus a 2026-08-06 three-worker field observation. +## [2026-08-12] ingest | knowledge-flush of 6 queued insights: 5 new pages, 1 dropped as an in-flight duplicate. New: testing/quality/signed-link-verification-assertions (assert a tokenized link through the verifier the receiver runs, with wrong-subject and wrong-key rejection arms in the same test and a token-stripped control on the live request — a signature is fixed-width for every key, so `?t=` presence is satisfied by a hardcoded placeholder; sources: AWS presigned-URL guide, docs.python.org hmac). backend/common/change-impact/cross-module-consumer-census (count production references to a task's new public symbols outside the defining module and outside its own tests, then classify by declared cross-module intent — knip's `ignoreExportsUsedInFile` and `includeEntryExports` encode the same internal-helper and entry-point populations a raw zero count cannot separate; measured 14 zero-consumer functions, exactly 1 a real gap). qa/process/defect-class-resweep-after-review (re-run each review finding's class search over the post-edit file including the lines the remediation just added; Yin et al. FSE'11 measured 14.8-24.4% of post-release fixes incorrect). infrastructure/containers/failing-pod-on-a-repo-synced-cluster (branch on pod phase before reading logs — `lastState.terminated` exitCode 137/OOMKilled is the runtime enforcing the limit and leaves no application log line; land the fix as a manifest PR because Argo CD does not sync live-cluster edits back to Git and self-heal reverts them). platforms/processes/cloud-cli-invocation-bounds (read the leaf subcommand's help, name region/profile/project/context on every invocation, cap list output with `--max-items` — the AWS CLI "retrieves all available items" by default). Dropped: the Python `.pyc`-cache mutation-harness insight, carried in better form by open PR #52, which additionally shows `-B` alone does not help when a stale `.pyc` already exists. Every cited URL was opened and quote-checked in this session. +## [2026-08-12] ingest | knowledge-flush of 4 queued insights (one arrived mid-flush) → 1 new page, 1 merged into an existing page, 1 folded into an in-flight PR, 1 dropped as an in-flight duplicate. New: databases/data-survey/catalog-statistics-as-current-state (a large table's newest rows: start the ctid tail scan at pg_relation_size/current_setting('block_size'), not at pg_class.relpages — relpages is "only an estimate ... updated by VACUUM, ANALYZE, and a few DDL commands", so appended blocks sit above it; LIMIT n on a ctid range returns the range's oldest end, so aggregate (DISTINCT/max/count) or narrow to the last blocks; read last_analyze/n_mod_since_analyze before citing pg_stats MCVs, and treat reltuples vs n_live_tup divergence as unmeasured inflow; append-mostly precondition, and TID range scans need PG14+). All eight cited URLs opened this session. Merged: testing/quality/tests-that-cannot-fail +1 mutation-outcome row / +1 edge case / +1 Instead-of (an assertion on the observable exists and still nothing reddens, because a later loop or error handler on the same path writes the same flag — assert from an input that leaves the other writers inert; reproduced: deleting the branch left 42 tests green, a `messages=[]` case turned it RED). Folded into open PR #73's testing/quality/default-values-under-test (mechanism-test fixtures must differ from the shipped default, or a "knob ignored, default always used" mutant is unkillable by any assertion). Dropped: the Python mutation-harness bytecode-cache insight — open PR #52 already carries it verbatim in backend/python/language/bytecode-cache-staleness (the -B/PYTHONDONTWRITEBYTECODE "governs writing only, purge __pycache__ once then keep -B" row). +## [2026-08-12] ingest | knowledge-flush of 3 queued insights → 2 ingested here, 1 folded into PR #64's branch. New: infrastructure/agent-orchestration/dispatching-after-a-completion-report (a worker's completion report settles the task, not the terminal — gate the next `worker-start` on `terminal wait --for tui-idle`, choose each settled dispatch's next owner (transfer/release/retain) before waiting again, read a failed start's receipt instead of re-running it, and retry through `--retry-of` with placement repeated, because 3 consecutive failures circuit-break the task into `failed`); platforms/tools/agent-permission-classifier-denials (read an auto-mode denial as one of four tiers — `permissions.deny`, `hard_deny`, `soft_deny`, allowed — because only `soft_deny` is cleared by an `allow` entry or by the user naming the exact action in their next message; write permission/`autoMode` config at user scope, since the classifier does not read `autoMode` from `.claude/settings.json` or `.claude/settings.local.json` at all, making a project-scoped write both denial-prone and inert). Sources re-verified this session: the auto-mode-config doc (four-tier precedence, the explicit-intent wording, the project-settings exclusion, `"$defaults"` splicing, `classifyAllShell`, the `Blocked by classifier` fixed reason, the `/permissions` retry path), anthropics/claude-code#58222 and #64128 (both closed as not planned), and the Orca CLI's own guide and `--help` output read live (`Wait for tui-idle before dispatching`; the `worker_done` next-owner rule; "After 3 consecutive failures on one task, the dispatch context circuit-breaks and the task is marked failed"; "The call exits 0 only for ready"; "--retry-of links the replacement attempt but does not inherit placement"). Both pages were drafted by an earlier flush session that never committed them — they existed only as untracked files in the shared checkout, so neither was on main nor in any open PR. +## [2026-08-12] ingest | knowledge-flush of 4 queued insights → 2 new pages, 1 merge, 1 candidate corrected. New: platforms/tools/plugin-mcp-server-registration (a plugin-bundled MCP server absent from `/mcp` is a registration fault first — `/reload-plugins` before any config edit, then `claude mcp list`/`--debug`, then a manual `initialize` run; both `.mcp.json` shapes load, so a shape rewrite that "fixes" it was really the reload; an unset `${VAR}` with no default is delivered to the server as literal text with a warning only). New: security/authn/retiring-a-replaced-auth-gate (after an auth cutover, census the retired session key's readers and writers — a key with readers and zero writers is the missed route; its "empty config → allow" fallback passes locally and refuses everyone in production, so parameterize the regression test on the setting that decides the fallback; deny default per ASVS 4.1.5). Merged: infrastructure/agent-orchestration/pane-delivery-confirmation +1 Instead-of (a send wrapper's success word means the keys reached the pane, not that the prompt was submitted — confirm from the pane) +2 edge cases (collapsed paste placeholder → send Enter as its own key event; a queued-then-picked-up wait pair is a confirmation) + the 2026-08-12 three-worker field observation. Corrected before ingest: the queued claim that a plugin stdio server inherits arbitrary exported shell variables so `env` can be dropped was not substantiated — the MCP reference stdio client passes only a fixed allowlist when `env` is absent, and the session's evidence measured a manual shell run rather than the harness spawn; the page keeps the `env` entry with a `:-` default instead. Sources live-verified: code.claude.com plugins-reference + mcp, modelcontextprotocol/typescript-sdk#216, OWASP ASVS V4.1.5, CWE-561, CWE-1188. +## [2026-08-13] ingest | knowledge-flush of 3 queued insights: 1 merged into an existing page, 2 retired. Merged into infrastructure/agent-orchestration/session-completion-gates: the receiving side of a completion gate — a worker on which the gate repeats at an instructed pause holds its phase and reports the gate-vs-prompt mismatch rather than advancing to a terminal value, with a phase-authorship table (worker writes `plan_ready`/`impl_done`, and `done` only after approval; `approved`/`merged` are the coordinator's review verdict). Reproduced at `fa89dc2`: `hooks/loop-gate.sh:55` accepts `done|approved|merged|failed|""` while `skills/orchestrate/scripts/status-update.sh:6` lists `impl_done`, and the three values that would silence the gate are exactly the three `skills/orchestrate/scripts/ready-set.sh:74` counts as dependency-satisfying — so a fabricated phase dispatches dependents against an unreviewed interface. Sourced to NIST SP 800-53 r5 AC-5 (separation of duties). Retired: the second copy of that insight (same trigger, same directive, harvested twice in one day) and the review-remediation defect-surface insight, folded into open PR #76's qa/process/defect-class-resweep-after-review rather than re-ingested here. +## [2026-08-13] ingest | knowledge-flush of 3 queued insights into 2 new pages (two insights share one page — same mechanism, one case). New: backend/java/jpa/not-null-check-and-lifecycle-callbacks — attribute the `PropertyValueException: not-null property references a null or transient value` by path shape rather than by the word "transient" (one hardcoded literal is shared by both throw sites in `Nullability`, and the only production occurrence in hibernate-orm; dots come solely from `buildPropertyPath` via `checkSubElementsNullability`, which recurses into `CompositeType` and into collections of composite elements), plus the ordering fact that `DefaultFlushEntityEventListener.scheduleUpdate` checks nullability *before* queueing `EntityUpdateAction` (5.6/6.6/7.0), so no `@PreUpdate` has run and a listener-based fix cannot work; INSERT-side "or transient" traced to `nullifyTransientReferencesIfNotAlready()`; and the silent-toggle case — `TypeSafeActivator` calls `setCheckNullability(false)` when validation mode is CALLBACK/AUTO and `hibernate.check_nullability` was never set ("Defaults to disabled if Bean Validation is present in the classpath and annotations are used, or enabled otherwise"), so a dependency change flips enforcement; measure via `isCheckNullability()` or a flush-and-expect test. New: databases/data-survey/audit-columns-as-update-evidence — an audit column is evidence only about the writer that sets it (Spring Data's `AuditingEntityListener.touchForUpdate` is `@PreUpdate`), so all-NULL `update_dt` cannot separate "never updated" from "every update failed pre-flush", and bulk JPQL leaves it untouched ("The effect of an `update` or `delete` statement is not reflected in the persistence context"); bound the claim, judge history on an independent axis, and require one positive control before acting on the absence. Sources live-verified this session against hibernate-orm 5.6/6.6/7.0 sources, the ValidationSettings/SessionFactoryOptions javadocs, the Hibernate query-language guide, and spring-data-jpa; field observation from PRD `manage.building_tenant_floor_info` (574 rows). +## [2026-08-13] ingest | knowledge-flush of 5 queued rows (4 unique insights; one row was a same-session Korean duplicate). New category security/incident-response +2 pages: verifying-assumed-security-agents (a documented EDR claim is verified on the host — vendor dir + process grep + `systemextensionsctl list` together, all-empty ⇒ "not installed" not "failed to detect"; derived from the SentinelOne-claimed / XMRig-4d8h incident, commands re-run this session) and process-identity-by-path-and-hash (a miner-suspect name matching a system daemon is judged by executable path + codesign, never name — genuine `/usr/libexec/sysmond` vs `~/.config/sysmond` XMRig; MITRE T1036.005). Merged: infrastructure/agent-orchestration/worktree-isolated-workers +1 edge case / +1 instead-of / +1 source (a repo-relative path to a gitignored coordinator-state dir resolves only in the main checkout — `git worktree add` never materializes ignored dirs; substitute absolute paths into worker briefs; local worktree reproduction) and debugging/methodology/reproduce-first +1 edge case / +1 field source (prod-only failure the client swallows: grep the service logs for the endpoint path before reading more code; ON CONFLICT schema-drift field case). Related links added both ways (worktree↔shared-run-state, reproduce-first↔logs-and-correlation, the two new pages to each other). +## [2026-08-14] ingest | knowledge-flush of 4 queued insights. New: qa/environments/headless-browser-bot-blocking (page shell renders but data APIs alone 4xx under a headless browser — probe with a desktop Chrome UA before diagnosing an outage; Chromium builds the headless default UA from kHeadlessProductName="HeadlessChrome"; field repro dabangapp 400→200), backend/common/integrations/estimate-derived-thresholds (absolute SL/TP-style triggers derived from a pre-execution estimate must be re-anchored ratio-preserving at the single confirmation-recording point — live incident: +2% design band collapsed to +0.06–0.42% via slippage; freqtrade anchors stoploss to entry after fill), testing/quality/history-dependent-checks-on-shallow-clones (guard git-history checks with --is-shallow-repository and skip honestly, or deepen fetch-depth per job; local repro: depth-1 boundary commit reported 306/306 files as added). Merged: qa/process/scope-purity-checks +Do#4 evidence-source-by-gate-lifetime table, +permanent-suite edge/instead-of rows, +Bazel test-encyclopedia hermeticity source (a permanent suite asserting `git status` reads someone else's in-progress tree — prove scope from the introducing commit's diff or skip). All URLs live-checked this session; two local reproductions (Chromium source constant via raw fetch; git 2.50.1 shallow-clone diff-filter). +## [2026-08-14] ingest | testing-quality-captured-log-message-assertions — assert captured log content via record.getMessage()/caplog.messages; record.message is pre-interpolated at capture, so `% record.args` re-application TypeErrors on parameterized calls (zero-arg calls mask the bug) +## [2026-08-14] ingest | debugging-methodology-probe-path-vs-operation-path — a passing precondition probe is evidence about the probe's path only; under refresh-token cookie auth a page-load login check diverges from direct API calls replaying a stored cookie jar — preflight the API's identity endpoint with the operation's own client +## [2026-08-14] ingest | knowledge-flush of 3 queued insights. New: backend/python/packaging/data-files-and-install-paths (NEW category packaging — `__file__`-relative data files break after non-editable install; declare package data + resolve via importlib.resources files()/as_file(); verify against a wheel installed into a scratch venv; setuptools + CPython docs verified, linkly v0.4.0 rc=4 reproduction). Merged: testing/data/test-data-and-isolation +1 Do row / +1 edge case (orchestrator-injected coordination env — unset in setup, pass tempdir state paths explicitly; leaked values also corrupt the live run's shared state; dev-loop #100 reproduction) with related↔shared-run-state; qa/document-verification/editing-a-gated-document +reflow breakage mode on the Substring anchor row, +platform-invisible-failure edge (bash ≤4.0 mid-test `[[ ]]` gap), +Instead-of row, measured 2026-08-14 + dev-loop PR #94/#102 field reproductions. Mechanism correction during verification: the harvested claim "macOS bash 3.2 matches a newline-split phrase in [[ ]]" was falsified by measurement — the wrap breaks the match on every platform; macOS only hides the failure (errexit compound-command gap, cross-linked to tests-that-cannot-fail / open PR #47). +## [2026-08-14] ingest | knowledge-flush of 3 queued insights, all new pages (no merge target had the mechanism). New: backend/common/change-impact/aggregation-layer-of-a-shared-helper (a plan naming "the N call sites" by line number can name sites in two different aggregation layers — a set-level SQL aggregate and a per-row projection reduced in application code grep identically; classify the layer, name one owner of the missing-value rule, and check the plan's grep-count criterion is reachable on the route you take; four routes disagree at the boundary and `SUM` over `[5, null]` returns a partial sum that reads as measured), testing/quality/generated-sql-property-assertions (locking "missing stays missing" on a rendered string CI never executes: refills split by POSITION, and the families are disjoint — top-level start/end anchors own the outside, a function-name check owns named inner wrappers, an aggregate-occurrence count owns the nameless inner CASE refill, a literal THEN NULL owns a guard rewritten to yield 0, a captured-alias binding owns dead-column swaps; build the coverage matrix and run rename+reflow controls), qa/document-verification/retiring-a-provisional-marker (converting `[추정]`/TBD to settled statements: split the marker's hits into body / checklist-citing-the-marker / round-history / convention-legend axes, rewrite the checklist rows whose evidence was the marker in the same commit, annotate history rather than rewording it, and report counts with the axis and command). Corrections made to the drafted candidates during verification: the PostgreSQL "All these functions ignore null values in their aggregated input" quote is scoped to Table 9.64 ordered-set aggregates, not `sum` (replaced with `sum`'s own "Computes the sum of the non-null input values"); the MDN reduce quote was a paraphrase presented as verbatim (replaced with the real sentence); and the candidate's central claim that top-level anchors catch every refill was refuted by local reproduction — inner-position refills leave both anchors matching, and a nameless inner CASE refill survives anchors AND name check, killed only by an aggregate-occurrence count. Sources live-verified this session (PostgreSQL functions-conditional/functions-aggregate, MySQL 8.4 comparison-operators, MDN Array.reduce, Stryker mutator catalogue, POSIX grep, Nygard ADR, RFC 7322); measurements on PostgreSQL 16.14 + 17.11 and Node v24.8.0; six-family × six-mutant coverage matrix reproduced locally with two behaviour-preserving controls green. ## [2026-08-17] ingest | Wiki-audit gap seeds G1-G6 (issue #38, 6 new pages). New categories: backend/common/architecture (sync-vs-async-integration — direct call vs queue vs event by consistency/latency/failure-isolation, AWS Prescriptive Guidance), backend/common/realtime (websocket-sse-lifecycle — auth at handshake, ping/pong dead-peer detection, SSE retry/Last-Event-ID reconnect vs WebSocket client-side backoff, backpressure bounding, shutdown draining; RFC 6455 + WHATWG SSE spec). Existing categories: backend/common/api-design +2 (cors-and-preflight — simple vs preflighted requests, wildcard-forbidden-with-credentials; api-versioning-and-breaking-changes — backward-compatible vs breaking change classification, Sunset header deprecation signal; MDN, Fetch spec, Stripe versioning/upgrades docs, RFC 8594), infrastructure/deploy +1 (feature-flag-lifecycle — release/experiment/ops/permissioning categories, toggle-as-inventory, removal-task-on-introduction; Fowler FeatureToggles), databases/operations +1 (data-backfill-migrations — batch/transaction sizing so a backfill doesn't hold row locks for its full duration, resumability, completion verification; strong_migrations, Retool Postgres migrations guide). All cited URLs live-fetched and verified this session. G1's other two categories.md sub-topics (module-boundaries-and-layering, event-driven-adoption-criteria) intentionally not written — narrower single-page scope per t38 plan D2; left as future ingest candidates on issue #38. G7-G10 remain open on the issue, out of this task's scope. +## [2026-08-17] ingest | knowledge-flush of 3 queued insights. New: backend/common/change-impact/inserting-a-guard-before-an-existing-side-effect (a plan can get the guard *decision* right while getting the *placement* wrong — trace the target's actual line order, pin the insertion before the first statement that mutates the checked state, and test the refusal path by asserting the side effect did not run; Eiffel DbC "preconditions on routine entry" + dev-loop t106 reproduction). Merged: infrastructure/agent-orchestration/shared-run-state +1 edge case / +1 Instead-of / +2 sources (coordinator re-prompt status reset is last-write-wins — reset before sending, never after; linkly iss0817 t60: worker plan_ready overwritten by a post-send pending re-seed, watcher deadlocked), infrastructure/agent-orchestration/worktree-isolated-workers +1 edge case / +1 Instead-of / +2 sources (a Bash-hook guard never runs on native Edit/Write calls — hooks doc: matchers filter on tool name; brief must say relative-paths-for-all-tools and the coordinator must git-status the protected tree pre-merge; linkly field observation: worker edited two main-checkout files via Edit under a worktree_escape guard, zero block/log). Related links added both ways (call-site-enumeration, guard-true-path-coverage ↔ new page; shared-run-state → pane-delivery-confirmation; worktree-isolated-workers → control-signals-vs-primary-artifacts). Hooks-matcher mechanism verified against code.claude.com/docs/en/hooks this session; other two field-tested with in-session reproductions. ## [2026-08-18] ingest | WebMCP-derived durable practices (2 new pages, 2 new categories). frontend/agent-interfaces/agent-facing-tool-surfaces (NEW category agent-interfaces: declarative form attributes toolname/tooldescription/toolparamdescription vs imperative document.modelContext.registerTool by action shape, reuse the human UI's handler as execute, label→schema derivation makes semantic forms the schema source, feature-detect + additive-layer because the API is a CG draft in Chrome 149–156 origin trial, :tool-form-active visibility, text budgets), security/agent-exposure/in-session-tool-exposure (NEW category agent-exposure: page content is attacker input to an in-session agent — design for hijacked-agent blast radius, consequence-class confirmation table with autosubmit omitted for state-changing tools, readOnlyHint/untrustedContentHint, exposedTo stays same-origin, server-side controls unchanged because inputSchema is advisory). All 5 cited URLs live-fetched this session; noted divergence verified against the live spec: Chrome docs mention requestUserInteraction() but the CG draft does not define it, so confirmation directives rely only on page-own UI mechanisms. Deliberately NOT ingested from the source material: adoption advocacy ("add WebMCP now") and traffic-share claims — single-browser experimental status fails the directive bar. +## [2026-08-18] ingest | Consolidated review of knowledge PRs #47–#110 (30 open PRs, 2026-08-06→08-17) into one reconciled state on post-#111 main. 53 new pages, 51 amended pages, 12 indexes updated. New categories: backend/common/ml (mape-aligned-point-prediction), backend/python/packaging (data-files-and-install-paths), security/incident-response (verifying-assumed-security-agents, process-identity-by-path-and-hash). Branches merged chronologically; conflicts union-resolved (related-id unions, max last_verified, newer routing framing + branch-new trigger clauses for index rows, edge-table row unions with one duplicate-substance row dropped). Each branch's own log entries spliced into date order above; every branch's added pages verified present in the final tree. The 30 source PRs are closed in favor of this consolidation (same protocol as the #17–#40 and #42–#43 reconciliations). diff --git a/tests/wiki-lint-prohibitions.bats b/tests/wiki-lint-prohibitions.bats index 14a2a3e..d39b4be 100644 --- a/tests/wiki-lint-prohibitions.bats +++ b/tests/wiki-lint-prohibitions.bats @@ -18,11 +18,11 @@ setup() { # --- normal: the real corpus is already compliant --------------------------- -@test "real wiki: exits 0 with 0 violations and 64 directive units" { +@test "real wiki: exits 0 with 0 violations and 71 directive units" { cd "$REPO_ROOT" || return 1 run node "$CHECKER" wiki [ "$status" -eq 0 ] - [[ "$output" == *"directives: 64"* ]] + [[ "$output" == *"directives: 71"* ]] [[ "$output" == *"violations: 0"* ]] } diff --git a/wiki/backend/common/api-design/error-responses.md b/wiki/backend/common/api-design/error-responses.md index 4526e85..7893206 100644 --- a/wiki/backend/common/api-design/error-responses.md +++ b/wiki/backend/common/api-design/error-responses.md @@ -9,7 +9,7 @@ sources: - https://www.rfc-editor.org/rfc/rfc9110 - https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Status last_verified: 2026-07-10 -related: [backend-common-errors-exception-handling] +related: [backend-common-errors-exception-handling, backend-common-errors-diagnostics-from-a-shared-code-path] --- # Choosing Status Codes and an Error Body Shape for an API diff --git a/wiki/backend/common/api-design/unenforced-declarations.md b/wiki/backend/common/api-design/unenforced-declarations.md index d0e56cd..c191368 100644 --- a/wiki/backend/common/api-design/unenforced-declarations.md +++ b/wiki/backend/common/api-design/unenforced-declarations.md @@ -9,7 +9,7 @@ sources: - https://kubernetes.io/blog/2023/04/24/openapi-v3-field-validation-ga/ - https://json-schema.org/draft/2020-12/json-schema-validation last_verified: 2026-08-05 -related: [security-input-validation-at-trust-boundaries, infrastructure-config-environment-config, backend-common-api-design-error-responses, qa-process-acceptance-criteria] +related: [security-input-validation-at-trust-boundaries, infrastructure-config-environment-config, backend-common-api-design-error-responses, qa-process-acceptance-criteria, backend-common-change-impact-widening-a-closed-value-table, platforms-processes-tool-diagnostics-without-a-failing-exit-code, backend-common-errors-diagnostics-from-a-shared-code-path] --- # Accepting a Declaration the System Does Not Enforce diff --git a/wiki/backend/common/change-impact/aggregation-layer-of-a-shared-helper.md b/wiki/backend/common/change-impact/aggregation-layer-of-a-shared-helper.md new file mode 100644 index 0000000..d59d785 --- /dev/null +++ b/wiki/backend/common/change-impact/aggregation-layer-of-a-shared-helper.md @@ -0,0 +1,119 @@ +--- +id: backend-common-change-impact-aggregation-layer-of-a-shared-helper +domain: backend +category: change-impact +applies_to: [general] +confidence: verified +sources: + - https://www.postgresql.org/docs/current/functions-aggregate.html + - https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/Array/reduce + - https://pubs.opengroup.org/onlinepubs/9799919799/utilities/grep.html +last_verified: 2026-08-14 +related: [backend-common-change-impact-call-site-enumeration, databases-schema-design-nullability-and-defaults, testing-quality-checks-that-cannot-pass, testing-quality-generated-sql-property-assertions] +--- + +# A Plan Unifying Call Sites That Sit in Two Different Aggregation Layers + +## When this applies + +A plan, task brief, or review comment tells you to unify or replace "the N call +sites" of a helper, and identifies those sites by line number rather than by what +the code does. One of the sites feeds a set-level aggregate in the database while +another returns a per-row value that application code reduces later. Also when a +"shared helper" change landed with one consumer adopting it while the other kept +its old semantics. + +Establishing that the site list is _complete_ → +[backend-common-change-impact-call-site-enumeration]. This page is about the +sites the list already names. + +## Do this + +1. **Open each named site and record which layer owns the reduction**, before + editing anything. Grep reports both as a call to the same helper; the owner of + the roll-up rule is what decides whether one helper can serve both: + +| What the site does | Who owns the roll-up | Consequence for a shared helper | +|--------------------|----------------------|---------------------------------| +| The helper's output sits inside a set-level aggregate (`SELECT SUM(helper(col))`) | The database's aggregate semantics | Changing the helper changes the rolled-up value | +| The helper's output is a per-row projection (`vacancyArea: helper(col)`) consumed by a later in-language reduce | The application function that reduces the rows | Changing the helper leaves the roll-up rule untouched — the reduce still decides the result | +| The helper's output is returned to a caller that neither aggregates nor reduces | The caller | Neither of the above rules applies; treat it as a third route | + +2. **Compare the two layers at the empty and all-missing boundary**, because that + is where they diverge and where a unification silently picks one convention. + PostgreSQL computes `sum` over "the non-null input values", and "except for + `count`, these functions return a null value when no rows are selected. In + particular, `sum` of no rows returns null, not zero as one might expect." A + seeded reduce has the opposite default — MDN: "if `initialValue` is provided but + the array is empty, the solo value will be returned without calling + `callbackFn`". A rule that treats any missing input as missing output agrees + with none of them: + +| Input | SQL `SUM(v)` | `reduce((a,b)=>a+(b??0), 0)` | `filter(non-null)` then reduce, `null` when empty | strict: any missing → missing | +|-------|--------------|------------------------------|---------------------------------------------------|-------------------------------| +| `[5, null]` | `5` | `5` | `5` | `null` | +| `[null, null]` | `NULL` | `0` | `null` | `null` | +| `[]` (no rows) | `NULL` | `0` | `null` | `null` | + +Measured 2026-08-14 (PostgreSQL 16.14 and 17.11; Node v24.8.0). No two columns +agree everywhere, so a plan that moves a rule from one column to another is +changing behaviour even when the helper's text is identical. The `[5, null]` row +is the dangerous one: three of the four routes return a **partial sum** that reads +downstream as a measured total, so a domain whose policy is the fourth column +cannot express its rule with a bare aggregate at all. + +3. **Check the plan's acceptance criterion against each route before adopting + it.** A criterion phrased as a grep count ("this helper appears at exactly two + sites") is reachable only on the route where both consumers call it. State + which route makes it reachable, or replace the criterion with one that holds on + the route you take — a criterion no route satisfies turns the task into + improvisation ([testing-quality-checks-that-cannot-pass]). + +4. **Name one owner of the rule and write it into the plan before editing.** Three + routes exist: push the reduce into SQL so one aggregate owns it; keep the reduce + and have the helper render only the per-row value it consumes; or give the rule + its own pure function and make every other route — the SQL included — a declared + _mirror_ of it. The third scales past two consumers, because the mirror + relationship is what a test can assert; the first two leave the rule wherever it + already was. Recording the choice is what stops the next round re-deriving it. + +5. **Assert that every consumer agrees, not that each one is individually + plausible.** Feed one fixture through all routes and require identical output, + including at the boundary. A per-consumer test passes while two consumers hold + different rules — agreement is the property the unification was for, and for the + SQL route it is asserted on the rendered expression + ([testing-quality-generated-sql-property-assertions]). + +6. **Re-run the enumeration after the edit and report the count with its + command**, since `grep -c` writes "only a count of selected lines" — a site + whose call spans two lines, or two calls on one line, moves the number without + moving the code. + +## Edge cases + +| Case | Then | +|------|------| +| The reduce is shared with an unrelated field (the same `sumCounts` also folds a parking count) | Its callers are a second enumeration pass — changing the reduce to match the aggregate changes every field it folds, so change the call site rather than the shared reduce | +| Only one consumer's boundary behaviour is specified by the policy document | Implement that one against the policy and record the other as an open decision in the plan; matching it to the specified one by symmetry invents a rule nobody approved | +| The per-row route is a paginated list and the aggregate route is a detail view | They can legitimately differ in cost but not in value; keep one owner of the rule and let the other read the same rendered expression | +| The helper is called from a raw SQL string as well as the builder | The string site is not statically enumerable by the builder's API — enumerate the string form too and record it in the plan | +| A single site is both: a per-row projection that a window function also aggregates | The database owns both; treat it as row one of the step-1 table and drop the application reduce | +| The plan names line numbers and the file has since moved | Re-derive the sites from the code shape, not the numbers, and update the plan — a stale line number points at a site that does something else now | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Adopt "replace the two call sites of this helper" because grep shows two hits | Open both and record which layer reduces the value | A per-row projection and a set-level aggregate are the same text and different operations; only one of them is changed by editing the helper | +| Land the helper at the SQL site and leave the list route on its old reduce | Choose the owning layer in the plan and migrate both consumers, or narrow the task to one consumer explicitly | A helper with one adopted consumer is the producer/consumer drift the unification existed to remove, and both sides stay green | +| Report the task done because the grep count matches the plan | Feed one fixture through every consumer and require identical output at the boundary | The count is satisfied by the text; agreement is the property, and a per-consumer test is green while two consumers still hold different rules | +| Read a `SUM` that returns a number as evidence the value was measured | Check whether any summand was missing before trusting the total | `sum` skips nulls, so `[5, null]` returns `5` — a partial sum reaches sorting and display as though it were complete | +| Wrap the aggregate in `COALESCE(..., 0)` so it matches the seeded reduce | Decide whether zero and absent mean the same thing for this field, and make both layers say so | `sum` of no rows is null by specification; refilling it makes "not measured" indistinguishable from "measured as zero" ([databases-schema-design-nullability-and-defaults]) | + +## Sources + +- https://www.postgresql.org/docs/current/functions-aggregate.html — `sum` "Computes the sum of the non-null input values"; "It should be noted that except for `count`, these functions return a null value when no rows are selected. In particular, `sum` of no rows returns null, not zero as one might expect". (The page's sentence "All these functions ignore null values in their aggregated input" belongs to Table 9.64, the ordered-set aggregates — not to `sum`; the per-function description above is the sourced statement for `sum`.) +- https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/Array/reduce — "If the array only has one element (regardless of position) and no `initialValue` is provided, or if `initialValue` is provided but the array is empty, the solo value will be returned without calling `callbackFn`"; a `TypeError` is "Thrown if the array contains no elements and `initialValue` is not provided" +- https://pubs.opengroup.org/onlinepubs/9799919799/utilities/grep.html — `-c`: "Write only a count of selected lines to standard output" — the count is over lines, not matches +- Measurement 2026-08-14 (PostgreSQL 16.14 and 17.11, both in Docker): over two all-NULL rows `sum(v)` → `NULL`; over zero rows → `NULL`; over `[5, NULL]` → `5`. Node v24.8.0 on the same values: `[null,null].reduce((a,b)=>a+(b??0),0)` → `0`, `[].reduce((a,b)=>a+b,0)` → `0`, `[].reduce((a,b)=>a+b)` → `TypeError: Reduce of empty array with no initial value`, and a filter-then-reduce returning `null` for an empty remainder → `null` — the table in step 2 +- Field incident 2026-08-14 (rtb-unified, NEWRTB-2786 building vacancy roll-up): a plan named two SQL call sites of a vacancy-area helper by line number. One was a real `SELECT … SUM(...)`; the list value came from a per-block projection reduced in TypeScript by `deriveDetailBlockRollup` → `sumCounts`, which filters nulls and returns null only when _all_ are null — the skip-missing column of step 2, while the policy wanted the strict column. Adopting the plan as written would have landed the helper at the aggregate while the list kept its reduce, and its "two call sites" acceptance grep was unreachable on that route. Verified in the worktree 2026-08-14: the shipped resolution took step 4's third route — a pure `sumVacancyAreaStrict` single-owns the rule (empty → null, any null → null), the sort SQL is a declared mirror (`CASE WHEN bool_or(v IS NULL) THEN NULL ELSE SUM(v) END`), all three consumers (detail, list, sort) are asserted to agree by `building-vacancy-path-parity.test.ts` (15 tests passing), and `sumCounts` was deliberately left in place for the parking-count axis whose missing semantics differ. The plan's line numbers had already drifted by the time of this check (`sumCounts` at `:1658`, not `:1644`), which is the last edge case above diff --git a/wiki/backend/common/change-impact/call-site-enumeration.md b/wiki/backend/common/change-impact/call-site-enumeration.md index 8524d68..9f308e1 100644 --- a/wiki/backend/common/change-impact/call-site-enumeration.md +++ b/wiki/backend/common/change-impact/call-site-enumeration.md @@ -9,7 +9,7 @@ sources: - https://docs.python.org/3/library/ast.html - https://peps.python.org/pep-0570/ last_verified: 2026-08-05 -related: [qa-process-regression-scope, backend-python-language-mutable-state-traps, testing-data-test-data-and-isolation] +related: [qa-process-regression-scope, backend-python-language-mutable-state-traps, testing-data-test-data-and-isolation, testing-quality-policy-at-several-return-sites, backend-common-change-impact-widening-a-closed-value-table, backend-common-change-impact-corpus-sweep-before-a-rejection-rule, backend-common-errors-diagnostics-from-a-shared-code-path, backend-common-change-impact-inserting-a-guard-before-an-existing-side-effect] --- # Enumerating Call Sites Before Changing a Callee's Contract diff --git a/wiki/backend/common/change-impact/corpus-sweep-before-a-rejection-rule.md b/wiki/backend/common/change-impact/corpus-sweep-before-a-rejection-rule.md new file mode 100644 index 0000000..837f9f7 --- /dev/null +++ b/wiki/backend/common/change-impact/corpus-sweep-before-a-rejection-rule.md @@ -0,0 +1,88 @@ +--- +id: backend-common-change-impact-corpus-sweep-before-a-rejection-rule +domain: backend +category: change-impact +applies_to: [general] +confidence: verified +sources: + - https://github.com/AriPerkkio/eslint-remote-tester + - https://github.com/rust-lang/rust-clippy/blob/master/lintcheck/README.md + - https://github.com/rust-lang/crater +last_verified: 2026-08-07 +related: [backend-common-change-impact-call-site-enumeration, backend-common-api-design-unenforced-declarations, testing-quality-guard-shape-vs-consequence, qa-process-regression-scope] +--- + +# Bounding a New Rejection Rule Against the Existing Corpus + +## When this applies + +You are adding a rule to a compiler, linter, parser, schema validator, or repo +gate that will start rejecting input the tool has been accepting silently, and +the existing corpus — repo sources, fixtures, shipped examples, downstream +configs — has to keep passing. Also when such a rule landed and turned red on +inputs nobody had classified as defective. + +Deciding whether unimplemented declarative input should reject, warn, or be +ignored at all → [backend-common-api-design-unenforced-declarations]. A shape +guard that is already red on one legitimate artifact → +[testing-quality-guard-shape-vs-consequence]. + +## Do this + +1. **Implement the accept/reject predicate once as a throwaway script, outside + the production tree.** Thirty to eighty lines that read one input and print a + verdict. It needs neither the real IR, the real diagnostic plumbing, nor the + real config surface — only the same decision. + +2. **Run it over every input the shipped rule will meet, and put two things in + the plan: the reject count and the full list of rejected paths.** The list is + the reviewable artifact; a bare count cannot be checked by anyone. + +3. **Require that list to equal the set already known to be defective before + writing production code.** An extra entry is a false positive and the rule's + boundary is wrong. A missing entry means the rule does not cover the case + that motivated it. Both verdicts are cheap while no production code exists. + +4. **State the enumeration method next to the count** — "148 sources (40 `.lnpl` + files + 108 triple-quoted inline programs under `tests/`)". An unstated basis + is what makes a partial index invisible + ([backend-common-change-impact-call-site-enumeration]). + +5. **Re-run the same script after the production rule lands and require the two + verdict sets to agree.** A divergence there is a wiring defect in the real + rule, not a rule-boundary question, and the sweep is what separates them. + +| Case | Do | +|------|----| +| The rule has several independent clauses | Sweep each clause as its own pass and report per-clause counts; one aggregate number hides both a clause that matches nothing and a clause that matches everything | +| Corpus inputs exist in more than one form (files on disk, strings embedded in tests, generated fixtures) | Enumerate each form as a separate pass with its own count — a single-form sweep is a partial index | +| The corpus is not yours (published configs, downstream repos, a package registry) | Sweep a sample with the same predicate, report the sample size, and ship the rule as a warning level first | +| The rejected set is legitimate and large | The change is a narrowing you cannot land at once: put the rule behind an opt-in strictness level and migrate the corpus under it | + +## Edge cases + +| Case | Then | +|------|------| +| Inputs are assembled at runtime (concatenation, `.replace()`, templating) rather than stored literally | A text sweep cannot index them — enumerate the assembly sites by hand, verify each one individually, and record them in the plan as the sweep's stated blind spot | +| The sweep rejects zero inputs | Treat it as unproven, not clean: feed it one input you know the rule must reject and require a reject before believing the zero ([testing-quality-checks-that-cannot-pass]) | +| The corpus contains the bug's own reproduction case | It is supposed to be rejected — list it in the plan as an expected reject instead of exempting it | +| The tool's own suite carries deliberately invalid inputs | Exclude negative-test fixtures by location and name the excluded directory next to the count | +| The throwaway predicate and the landed rule disagree on an input | Re-run both on that input and fix whichever one contradicts the accepted list in the plan; the plan's list is the decision of record | +| The corpus is large enough that a full pass is slow | Keep a cheap textual prefilter and run the real predicate only on what it matched — the reported set is still the full-corpus answer | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Implement the rule in the production tree and let the test suite report what breaks | Sweep with a throwaway predicate and settle the boundary before writing production code | A red suite does not say whether the rule is wrong or the input was always defective — and once the code exists, the cheap resolution is to loosen the rule | +| Cite "the suite is still green" as evidence the rule is non-breaking | Cite the reject count and the rejected-path list from a full-corpus sweep | The suite exercises the inputs it happens to carry; the corpus is the set the rule will actually meet | +| Sample a few representative inputs | Run every input through the predicate | Boundaries fail on the unusual input, which is the one sampling drops | +| Add an exemption for the legitimate input the rule rejected | Narrow the rule's condition until the reject list matches the known-defective set | An exemption list is the rule conceding its boundary is wrong, one input at a time ([testing-quality-guard-shape-vs-consequence]) | + +## Sources + +- https://github.com/AriPerkkio/eslint-remote-tester — the questions a corpus run answers are "Does the rule report the intended patterns? Does the rule falsely mark valid patterns as errors?", and "the AST of Javascript and Typescript can cause very unexpected results it is not enough to test the rule only against unit tests and a small amount of repositories"; its comparison mode reports "the exact changes in ESLint reports their code changes introduced" +- https://github.com/rust-lang/rust-clippy/blob/master/lintcheck/README.md — lintcheck "Runs Clippy on a fixed set of crates read from `lintcheck/lintcheck_crates.toml` and saves logs of the lint warnings into the repo. We can then check the diff and spot new or disappearing warnings" — the recorded, diffed verdict list is the deliverable, not the count +- https://github.com/rust-lang/crater — "Crater is a tool to run experiments across parts of the Rust ecosystem. Its primary purpose is to detect regressions in the Rust compiler, and it does this by building a large number of crates, running their test suites and comparing the results between two versions of the Rust compiler" — the same measurement at ecosystem scale, run before a potentially breaking change lands +- Field evidence 2026-08-07 (`linkly` #53, Python): three draft rejection rules for a DSL compiler were implemented first as a throwaway `sweep.py` and run over 148 sources (40 `.lnpl` files plus 108 triple-quoted inline programs in tests) → 2 rejects, both the QA probes the issue had named. The sweep caught a false positive in the first draft at plan time: it rejected the legitimate case where a guard owns its own block. Fixtures assembled with `.replace()` were invisible to the text sweep, so those 5 sites were confirmed by hand — the blind spot from the edge-case table, observed rather than hypothesized +- The pre-implementation ordering (throwaway predicate before production code) is the field-tested refinement of this page; the corpus-sweep-and-diff method itself is the practice the three tools above implement diff --git a/wiki/backend/common/change-impact/cross-module-consumer-census.md b/wiki/backend/common/change-impact/cross-module-consumer-census.md new file mode 100644 index 0000000..61d92e3 --- /dev/null +++ b/wiki/backend/common/change-impact/cross-module-consumer-census.md @@ -0,0 +1,87 @@ +--- +id: backend-common-change-impact-cross-module-consumer-census +domain: backend +category: change-impact +applies_to: [general] +confidence: verified +sources: + - https://knip.dev/guides/handling-issues + - https://knip.dev/reference/configuration +last_verified: 2026-08-11 +related: [backend-common-change-impact-call-site-enumeration, testing-quality-tests-that-cannot-fail, infrastructure-agent-orchestration-worktree-isolated-workers] +--- + +# Counting the Production Consumers of a Symbol Your Task Just Added + +## When this applies + +Your task added a public function, endpoint, export, or hook whose consumer +belongs to a *different* task — parallel work split by file ownership, a backend +change whose UI wiring is another ticket, an agent worker producing a seam for a +sibling worker. You are deciding whether the task is done. Also when a feature is +typed, tested, reviewed, and merged, and changes nothing at runtime. + +Enumerating the callers of a symbol whose contract you are *changing* → +[backend-common-change-impact-call-site-enumeration]. + +## Do this + +1. **Take the new public symbols from the diff, not from memory**: names added by + `git diff origin/main...HEAD` in the files you own. That list is the census's + subject. + +2. **Count references outside the defining module, excluding the symbol's own + tests.** Grep the bare name across the repo, drop hits in the defining file + and in test paths, and record the remaining file list per symbol. The tests + are what make an orphan look alive — a symbol its own test calls has a + non-zero reference count and no production reachability + ([testing-quality-tests-that-cannot-fail]). + +3. **Count references, not calls.** Search the bare name as well as `name(`: + seam injection (`estimate_fn=estimate.estimate`), callback registration, + decorator tables, and registry entries all pass the symbol as a value, and a + paren-anchored search reports those wirings as dead. + +4. **Classify each zero-consumer symbol by declared intent, and treat only the + intent-bearing ones as defects.** The defect is a symbol whose docstring, plan, + task brief, or PR body states that another module consumes it. A zero count on + its own is the normal shape of a same-module helper. Unused-export tooling + encodes the same three populations: knip's `ignoreExportsUsedInFile` exists + because "In files with multiple exports, some of them might be used only + internally", and its `includeEntryExports` exists because "By default, Knip + does not report unused exports in entry files" — internal helpers and entry + points are the two populations a raw count cannot separate from real orphans. + +5. **Run the census at each task's review, not at integration.** At integration + every task is already approved, so the missing wiring has no owner; at review + the owning session is still open. + +6. **When the consumer task already finished, open a follow-up task naming the + file and the insertion point** — that task is the only thing standing between a + complete implementation and dead code. + +## Edge cases + +| Case | Then | +|------|------| +| The symbol is dispatched dynamically (`getattr`, a name in YAML/config, a route string) | Search the string form too, and record in the task that this class of site is not statically enumerable | +| The symbol is itself an entry point (CLI command, HTTP route handler, hook) | Its consumer is a registration, not a call — assert the registration file lists it (route table, plugin manifest, entry map) instead of counting references | +| A working language server exists | Use find-references for the count and state that as the method; text search stays the fallback for aliased re-exports | +| The consumer lives in another repository or a published package | The census cannot see it: record the consuming repo and the version that will adopt it, and keep the symbol out of the defect list | +| The symbol is re-exported through a package `__init__` or facade | Count references to the re-exported name as well, or every facade consumer reads as zero | +| The census returns zero for every new symbol | Suspect the search, not the code: confirm the pattern matches one symbol you know is wired before reading any zero as a finding | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Report every symbol with no reference outside its defining file as dead code | Filter that list by declared cross-module intent, and report those | Internal helpers dominate the raw list; measured on one module, 14 of 14 zero-consumer functions were reported and exactly 1 was a defect, so the unfiltered list buries the finding | +| Search `name(` to find consumers | Search the bare name as well | Seam injection and callback registration pass the symbol as a value and never write the paren | +| Treat "type-check, tests, and CI all green" as proof the wiring landed | Run the consumer census before calling the task done | Nothing in a type system or a test suite requires a new public symbol to have a production caller | +| Defer the census to the integration branch | Run it in each task's review | After every task is approved, the missing 3 lines belong to nobody | + +## Sources + +- https://knip.dev/guides/handling-issues — a surprising unused-export report "is usually a real finding or a configuration gap, not a false positive to silence"; before deleting, check whether the export is in an entry file, re-exported from another entry point, or tagged for external use — the report is a candidate list that intent resolves +- https://knip.dev/reference/configuration — `ignoreExportsUsedInFile`: "In files with multiple exports, some of them might be used only internally. If these exports should not be reported, there is a `ignoreExportsUsedInFile` option available"; `includeEntryExports`: "By default, Knip does not report unused exports in entry files" +- Field measurement 2026-08-11 (Python module set, 6 tasks split across parallel workers by file ownership): a census of every public function counted cross-module production references; 14 came back zero. Exactly one was a real gap — a URL-building helper whose docstring named its consumer ("the caller puts this in the Slack body") and which no producer of that message ever called, leaving the notification's approval link unsigned. The other 13 were same-module helpers. All 6 tasks had passed review, 402 tests were green, and the merge had no conflicts diff --git a/wiki/backend/common/change-impact/inserting-a-guard-before-an-existing-side-effect.md b/wiki/backend/common/change-impact/inserting-a-guard-before-an-existing-side-effect.md new file mode 100644 index 0000000..9f7f659 --- /dev/null +++ b/wiki/backend/common/change-impact/inserting-a-guard-before-an-existing-side-effect.md @@ -0,0 +1,67 @@ +--- +id: backend-common-change-impact-inserting-a-guard-before-an-existing-side-effect +domain: backend +category: change-impact +applies_to: [general, shell] +confidence: field-tested +sources: + - https://www.eiffel.org/doc/eiffel/ET-_Design_by_Contract_(tm)%2C_Assertions_and_Exceptions + - "Field reproduction (dev-loop t106, 2026-08-17): status-update.sh:19 unconditionally created a default status file for every phase; the plan's proposed guard placement sat after that line, so the guard could never observe the missing file it checked for" +last_verified: 2026-08-17 +related: [backend-common-change-impact-call-site-enumeration, qa-exploratory-guard-true-path-coverage] +--- + +# A New Precondition Guard Inserted Into Code That Already Mutates the Checked State + +## When this applies + +You are implementing, adopting, or auditing a planned change that adds a +precondition guard ("refuse when the file is missing", "abort when X is absent") +to an existing script or function — and the plan describes the insertion point in +prose ("after the extra is built", "at the start of phase 2") rather than by the +target's actual lines. Especially when an earlier line in the target already has +an unconditional default/auto-create side effect. + +## Do this + +1. **Trace the target's actual execution order before placing the guard.** Read + the file and pin the insertion point between named line numbers, not by the + plan's phase prose. A plan can get the *decision* right while getting the + *placement* wrong: "after X is built" reads as a natural phase description + but ignores an unrelated earlier line that already mutated the state. +2. **Confirm the guard sits textually before the first statement that can mutate + the state it checks.** An earlier `[ -f "$file" ] || create-default` makes a + later "refuse if file missing" guard dead code — the state it tests for can + no longer occur. This mirrors precondition doctrine: a precondition is + checked *on entry*, before any of the body runs (Eiffel monitors + preconditions "on routine entry"). +3. **When the planned placement lands after such a side effect, surface the line + numbers instead of silently re-planning.** In orchestrated or reviewed work, + report "guard at plan-position L would run after side effect at line N" and + let the plan owner move it — the decision itself is unchanged; only the + placement moves. +4. **Test the guard's negative path against the side effect, not just the exit + code.** Assert both the refusal (exit code/error) *and* that the earlier side + effect did not run (the default file was not created). A guard placed after + the auto-create can return the right code with the wrong state. + +## Edge cases + +| Case | Then | +|------|------| +| The side effect lives in a sourced helper or a function called earlier | Trace through the call, not just the top-level lines — "textually before" means before in execution order | +| The side effect is itself conditional (runs only for some inputs/phases) | Decide placement per path: the guard must precede it on every path where both can apply | +| The plan's prose position and the correct line position coincide | Still record the line numbers in the plan/PR — the next inserted line inherits the same ambiguity | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Implement the guard at the position the plan's prose describes | Read the target file and pin the insertion between named lines | Prose phases do not map one-to-one onto line order; an earlier unconditional line can precede "phase start" | +| Test the new guard by exit code alone | Also assert the pre-existing side effect did not fire | The refusal can be correct while the state the guard was meant to protect is already mutated | +| Quietly move the guard yourself in delegated/orchestrated work | Escalate with both line numbers and a proposed snippet | The plan owner may know a reason for the ordering; a silent rewrite hides the defect class from review | + +## Sources + +- https://www.eiffel.org/doc/eiffel/ET-_Design_by_Contract_(tm)%2C_Assertions_and_Exceptions — preconditions are monitored "on routine entry", i.e. before the body executes; a check placed after a mutation is not a precondition +- Field reproduction (dev-loop t106, 2026-08-17): `status-update.sh:19` unconditionally created `{"task":"…"}` for any phase; the originally planned rework-guard placement sat after it. Escalated with line numbers; the revised plan moved the guard between lines 18–19, and the BATS case "rework with no status file → exit 4, **no file created**" passed only with that placement diff --git a/wiki/backend/common/change-impact/widening-a-closed-value-table.md b/wiki/backend/common/change-impact/widening-a-closed-value-table.md new file mode 100644 index 0000000..baaf2b4 --- /dev/null +++ b/wiki/backend/common/change-impact/widening-a-closed-value-table.md @@ -0,0 +1,79 @@ +--- +id: backend-common-change-impact-widening-a-closed-value-table +domain: backend +category: change-impact +applies_to: [general] +confidence: verified +sources: + - https://refactoring.com/catalog/replaceMagicLiteral.html + - https://pragprog.com/tips/ +last_verified: 2026-08-06 +related: [backend-common-change-impact-call-site-enumeration, backend-common-api-design-unenforced-declarations] +--- + +# Widening a Closed Value Table Whose Consumers Inlined It + +## When this applies + +You are adding an entry to a closed table that maps a name to a magnitude or a +code — duration units, status codes, currency exponents, retry tiers, severity +levels — and the table exists as a named constant in one module. Also when a +newly added entry is accepted at one layer and rejected, ignored, or +mis-converted at another, so the symptom is a divergence between two paths +rather than a parse error. + +## Do this + +1. **Enumerate by the table's values, not only by its name.** Grep a + distinctive magnitude from the table (`60000`, `86400`, `4290`) and a + distinctive member string (`"ms"`), across the whole repo including tests, + fixtures, generators, and any second-language backend. The name grep lists + the sites that import the table; the value grep lists the sites that copied + it, and only the second set is where widening breaks. +2. **Treat the gap between the two counts as the work item**, and state both in + the plan: "1 site imports `DURATION_UNITS`, 7 mention `60000`" is checkable + and shows the scope; "the units table has 1 consumer" hides it. +3. **Classify every value hit before editing:** + +| Value hit is | Do | +|--------------|----| +| The canonical table's own definition | Nothing — this is the site the others should read | +| An inlined copy of the pairs (`(("ms",1),("s",1000),("m",60000))`) | Replace the literal with a read of the canonical table | +| Bare arithmetic on one member (`value % 60000`, `ms // 86400000`) | Replace the literal with a lookup into the table; it is a copy of one row | +| A second named table over the same vocabulary in the same module | Derive one from the other so a single edit widens both | +| A copy in another language, a generated artifact, or a backend that re-implements the conversion | Generate it from the table, or add a conformance test asserting both accept every member | + +4. **Fix every copy to read the single table before adding the new entry**, so + the entry lands in one place. Widening first and reconciling after means the + new member exists in the table while each copy silently defines the old, + narrower vocabulary. +5. **Add a test that drives every consumer with every member of the table**, + parameterized over the table itself. It fails on the next widening if a new + copy has appeared, which is the only check that survives the next author. + +## Edge cases + +| Case | Then | +|------|------| +| The inlined copies are already narrower than the canonical table | The divergence predates your change — the table's later members are already unreachable through those paths. Fix them as part of this change and note the pre-existing gap, or your widening gets blamed for it | +| The magnitude is not distinctive (`1`, `60`, `1000`) | Grep the member name string (`"ms"`, `"USD"`) and the suffix form instead; a bare `1000` returns every unrelated site and the enumeration stops being readable | +| The same vocabulary is expressed in two units across a boundary (seconds inside, milliseconds at the edge) | Grep both magnitudes (`60` and `60000`) — a copy converted at the boundary matches neither the table's values nor its name | +| Members of the table are persisted (stored in rows, serialized into messages, written into config already deployed) | Widening is a migration, not an edit: old readers must keep accepting stored values, and the new member cannot be written until every reader ships | +| The table is consumed by a caller you do not own (published package, other repo, plugin API) | Widening is a versioned release for them; enumerate only what you own, and treat the new member as unsupported until their version pins forward | +| A copy exists only in a test fixture | It still diverges — a fixture pinned to the narrow vocabulary keeps passing while the widened path is never exercised, so the suite reports green on an unmigrated consumer | +| The lookup is a `switch`/`if` chain over member names rather than a value copy | The value grep misses it — grep the member names as well, since the chain encodes the same table as control flow | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Grep the constant's name, report N consumers, and scope the change from that | Grep the table's values and member strings too, and scope from the union | Consumers routinely inline the pairs instead of importing the constant, so the name grep counts imports and misses copies — the copies are the sites that break | +| Add the entry to the canonical table and run the suite | Reconcile the copies first, then widen | A green suite means the consumers the tests reach accepted the entry; the copied ones were never asked | +| Leave one inlined copy because it is a hot path or avoids an import cycle | Have that path read the table once at import and keep the local binding | The copy is not cheaper than a module-level lookup, and it is the site that silently defines a different vocabulary | +| Extend a second same-vocabulary table alongside the first to keep both callers happy | Derive the second from the first in the same module | Two canonical-looking tables make the next author's name grep authoritative and wrong | + +## Sources + +- https://refactoring.com/catalog/replaceMagicLiteral.html — *Replace Magic Literal*, alias "Replace Magic Number with Symbolic Constant": a literal with a particular meaning becomes a named constant. The refactoring exists because the inlined literal is the default state of such values, which is what makes the value the reliable search handle +- https://pragprog.com/tips/ — Tip 15, DRY: "Every piece of knowledge must have a single, unambiguous, authoritative representation within a system." A copied value table is a second representation, and widening one representation is what produces the divergence +- Local reproduction 2026-08-06 (`linkly`, Python, macOS): `grep -rn "DURATION_UNITS" impl/lnpl/*.py` returns **1** hit (the definition); `grep -rn "60000" impl/lnpl/*.py` returns **7** across four files — a second named table (`DURATION_UNIT_MS`, `lexer.py:23`), three independently inlined `(("ms",1),("s",1000),("m",60000))` tuples (`condition.py:353`, `interp.py:1020`, `backend.py:446`), and two bare-literal arithmetic sites (`condition.py:269-270`). The predicted divergence was already present: the canonical map carries `h` and `d`, while all three inlined copies stop at `m`, so those paths cannot convert a unit the lexer accepts diff --git a/wiki/backend/common/errors/diagnostics-from-a-shared-code-path.md b/wiki/backend/common/errors/diagnostics-from-a-shared-code-path.md new file mode 100644 index 0000000..effa114 --- /dev/null +++ b/wiki/backend/common/errors/diagnostics-from-a-shared-code-path.md @@ -0,0 +1,86 @@ +--- +id: backend-common-errors-diagnostics-from-a-shared-code-path +domain: backend +category: errors +applies_to: [general] +confidence: verified +sources: + - https://rustc-dev-guide.rust-lang.org/diagnostics.html + - https://doc.rust-lang.org/stable/nightly-rustc/rustc_errors/enum.Applicability.html + - https://www.nngroup.com/articles/error-message-guidelines/ +last_verified: 2026-08-09 +related: [backend-common-api-design-error-responses, backend-common-api-design-unenforced-declarations, backend-common-change-impact-call-site-enumeration, debugging-signals-reading-error-messages] +--- + +# A Rejection Message Emitted From a Code Path Two Constructs Share + +## When this applies + +You are writing or reviewing a rejection/validation message produced by one +function that more than one caller reaches — a reference checker shared by a +guard and an assignment, a validator shared by request and response, a policy +check shared by two config blocks — and the message names one caller's construct +in literal text, or tells the author what to write instead. Also when a user +reports a rejection whose wording names a construct they did not write. + +## Do this + +1. **Make the message's subject a parameter the caller supplies.** The shared + function formats; each call site passes its own construct name. A construct + name written literally inside a shared body is correct for exactly one caller + and silently wrong for every other one. +2. **Verify each part of the message against the path that emits it**, because + the parts fail with different visibility: + +| Part the message carries | Verify by | +|--------------------------|-----------| +| The subject — the construct being rejected | Reading one emitted message per call path; a wrong subject is visible on first read | +| A named repair ("use `input.` instead") | Executing the repair as written from each call path and requiring the result to be accepted | +| A pointer to another rule or section | Confirming that rule admits this caller's construct at all | + +3. **Treat the repair as the part most likely to be wrong, and check it per + path.** A repair is advice the author will follow literally; when it is + illegal on one of the emitting paths, following it lands the author in a + second, unrelated rejection. NN/g's rule is that the message must describe a + solution sufficient to fix the problem — on the path the reader is on. +4. **Rate the repair by whether it holds on every path that can emit it.** + Present it as *the* fix (and allow any auto-apply tooling to use it) only when + it is valid on all of them; otherwise branch it. `rustc` encodes the same + distinction as `Applicability` — `MachineApplicable` for a suggestion that can + be applied mechanically, `MaybeIncorrect` for one that "may or may not be a + good one" — and instructs authors to "be conservative when choosing the level". +5. **Assert the message once per emitting path, not once per message.** A single + test on one caller leaves the other caller's subject and repair unasserted, so + parameterizing the subject and breaking the other path's advice both stay + green ([backend-common-change-impact-call-site-enumeration] enumerates the + paths). +6. **When the repair differs by path, branch on the parameter that already + distinguishes them** — pass the repair alongside the subject, so each caller + states the fix that is legal for it. + +## Edge cases + +| Case | Then | +|------|------| +| The two callers reject for the same reason but repair differently | Pass the repair text as a second caller-supplied parameter; keep one rejection rule and two suggestions | +| The suggested form is legal at parse time but rejected by a later rule on this path | The advice is still wrong — run it end to end on that path, not just past the check that emitted it | +| Only one caller exists today | Parameterize the subject anyway when the function is named for the *check* rather than the construct; the second caller is what makes the literal wrong, and it arrives without touching this file | +| The message is localized or templated | Pass the subject as a named placeholder argument rather than concatenating it, so translators receive a slot instead of a sentence fragment | +| The shared function genuinely cannot know the subject | Have callers pass a context value carrying subject and repair together, so a new caller cannot compile without supplying both | +| A repair is valid everywhere except one rarely reached path | Branch it — a suggestion that is wrong on one path is `MaybeIncorrect` for all of them, and rating it that way costs the reader on every path | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Write the construct's name into the shared function's message text | Take the subject as a parameter each call site fills | The literal is right for the caller you had in mind and misnames every other one | +| Fix the misnamed subject and ship | Also execute the repair from each emitting path and require acceptance | The subject is checked by reading; the repair is only checked by running, so it is the part that stays wrong | +| Soften the repair into something true on every path ("check the syntax") | Branch the repair on the parameter that distinguishes the callers | A repair that carries no action returns the reader to guessing, which is what the message existed to prevent | +| Assert the message text in one test and call the wording covered | Assert subject and repair once per emitting path | One assertion cannot distinguish "both paths right" from "one path never exercised" | + +## Sources + +- https://rustc-dev-guide.rust-lang.org/diagnostics.html — suggestions carry a confidence level and "Be conservative when choosing the level"; `MachineApplicable` = "Can be applied mechanically", `MaybeIncorrect` = "Cannot be applied mechanically because the suggestion may or may not be a good one", `Unspecified` = "we don't know which of the above cases it falls into" +- https://doc.rust-lang.org/stable/nightly-rustc/rustc_errors/enum.Applicability.html — the enum tools read to decide whether a suggestion is auto-applied or shown for review +- https://www.nngroup.com/articles/error-message-guidelines/ — an error message offers constructive advice: the described solution must be sufficient for the user to fix the problem +- Field incident 2026-08-09 (`linkly`, `impl/lnpl/lower.py`): `_Scope.check_reference` is called from both the guard path and the assignment path and hardcoded "guard condition" into three messages. The `set`-target rejection additionally advised writing `input.`, which `_derive_assignment` rejects for `set` targets by a separate rule — an author following the advice hit a second, unrelated rejection. After threading subject/target through as parameters the suite went 1864 → 1872 (the 8 new per-path assertions, no other change), and an independent audit exercised each branch as its own mutation diff --git a/wiki/backend/common/integrations/consumer-required-fields.md b/wiki/backend/common/integrations/consumer-required-fields.md new file mode 100644 index 0000000..be25f37 --- /dev/null +++ b/wiki/backend/common/integrations/consumer-required-fields.md @@ -0,0 +1,97 @@ +--- +id: backend-common-integrations-consumer-required-fields +domain: backend +category: integrations +applies_to: [general] +confidence: verified +sources: + - https://docs.pact.io/ + - https://json-schema.org/understanding-json-schema/reference/object +last_verified: 2026-08-10 +related: + [ + backend-common-change-impact-call-site-enumeration, + backend-common-integrations-externally-owned-defaults, + backend-common-llm-completion-response-validation, + testing-quality-tests-that-cannot-fail, + testing-quality-minimum-case-set, + ] +--- + +# Building a Payload for a Consumer Whose Required Fields You Inferred + +## When this applies + +You are writing an adapter that maps one module's records (a selection query, a +repository row, a scraped item) into the shape a second module consumes — a +scoring engine, a plugin, an external API client — and you took that shape from +the consumer's docstring, README example, or a sample payload. Also when such an +adapter runs end to end with no error and the downstream numbers look low, and +when deciding which mapped fields need an assertion of their own. + +Enumerating call sites when *you* own the callee → +[backend-common-change-impact-call-site-enumeration]. + +## Do this + +1. **Call the real consumer once with a mapped record before writing the rest of + the adapter.** An example payload is an instance, not a specification: in JSON + Schema terms "the properties defined by the `properties` keyword are not + required" unless a `required` list says so, so a sample cannot tell you which + keys the consumer depends on. Pact makes the same distinction — "unlike a + schema or specification (eg. OAS), which is a static artefact", a contract is + "enforced by executing a collection of test cases", and a test double is + trusted only when its calls "return the same results as a call to the real + application would". + +2. **Split the fields the consumer reads into loud and silent by how it reads + them**, and drive the split from the consumer's source, not its prose: + +| How the consumer reads the field | Class | Missing-field symptom | +|---|---|---| +| Presence check or direct subscript (`if k not in d`, `d[k]`, required-field validation) | Loud | Raises on the first record; the run stops | +| Read with a default (`d.get(k, …)`, `??`, `COALESCE`, optional destructuring) | Silent | No error; the computed value is wrong for every record | +| Used only for logging or display | Cosmetic | No error; output text is thinner | + +3. **Let the loud fields be found by one end-to-end run, and write an explicit + assertion for every silent field.** The loud ones announce themselves; the + silent ones cannot fail any test that only checks "the pipeline completed" + ([testing-quality-tests-that-cannot-fail]). + +4. **Make each silent-field assertion compare two runs of the real consumer — + one with the field populated, one with it dropped — and require the outputs + to differ.** Asserting only "the key is present in the mapped dict" passes on + a key the consumer never reads and on a value that never reaches the formula. + +5. **Record the discovered required-field list next to the mapping code**, with + the consumer version or commit you probed. The list is the adapter's real + contract, and step 1 is re-run against it when the consumer is upgraded + ([backend-common-integrations-externally-owned-defaults]). + +## Edge cases + +| Case | Then | +|------|------| +| The consumer is expensive or has side effects | Probe it once in a test with a single record and cache the field list — the probe is required even then, and the cache is what keeps it cheap | +| The consumer accepts the record but ignores unknown keys | The unknown key is not evidence of anything — keep the step-4 two-run comparison as the only proof a field is consumed | +| The consumer validates with a machine-readable schema (pydantic model, JSON Schema, protobuf) | Read `required`/non-default fields from the schema instead of probing, and still run step 4 for the defaulted ones | +| A silent field's absence changes the output by less than the assertion's tolerance | Choose the probe record that maximizes the field's contribution (longest text, largest count) so the two runs separate | +| The consumer defaults a missing field to a neutral value that looks plausible | Treat it as silent, not absent — a plausible default is what lets the defect reach production | +| Several silent fields feed one output number | Drop them one at a time; dropping all of them at once cannot attribute the difference | +| The consumer is non-deterministic (LLM scorer, sampling, clock-dependent) | Pin the seed/temperature or stub the non-deterministic part before comparing — otherwise the two runs differ for every field, including ones the consumer never reads, and step 4 passes vacuously | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Map fields from the consumer's docstring schema and move on | Call the consumer with one mapped record first, then map the rest | The docstring lists the shape, not which keys are load-bearing; the run tells you | +| Treat "the pipeline finished with no exception" as evidence the mapping is complete | Assert each silent field changes the output when dropped | The loud fields are the only ones an exception-free run proves | +| Assert `"desc" in payload` for each field the consumer reads | Run the consumer twice, with and without the field, and assert the outputs differ | Key presence is satisfied by a key the consumer ignores or reads from a different name | +| Add a default in the adapter for a field you are unsure the consumer needs | Probe first, then map the field or leave it out deliberately | A speculative default converts a loud failure into a silent one — the class that survives to production | + +## Sources + +- https://docs.pact.io/ — "The contract is generated during the execution of the automated consumer tests"; contract tests "check that all the calls to your test doubles return the same results as a call to the real application would"; and "unlike a schema or specification (eg. OAS), which is a static artefact that describes all possible states of a resource, a Pact contract is enforced by executing a collection of test cases, each of which describes a single concrete request/response pair" — the basis for step 1 preferring an execution over the documented shape +- https://json-schema.org/understanding-json-schema/reference/object — "By default, the properties defined by the `properties` keyword are not required"; an example payload therefore carries no required/optional distinction +- Reproduction 2026-08-10 (Python 3): a consumer that reads `assignee_id` with `if "assignee_id" not in item: raise ValueError` and computes `1.0 + 0.5 * len(item.get("desc", ""))`. Dropping `assignee_id` raised on the first record; dropping `desc` raised nothing and moved the returned score from 21.0 (`desc` of length 40) to 1.0 — same adapter output, one failure visible in the first run and one visible only to an assertion +- Field measurement 2026-08-10 (manday estimation engine, `manday-sp/engine.py`): the two read sites are `check_assignee_ids(items)` at line 511, whose contract is "키 존재 + 값 형식" (key presence, not just value shape), and `d = it.get("desc") or ""` at line 399 — loud and silent respectively. The split is a property of the read site, not of the field, so record the file:line (a sibling copy of the same scorer reads `it["desc"]` by subscript, which makes the same field loud). A mapping built from the engine's docstring schema produced `ValueError` on 100% of records for the missing `assignee_id` key, while the missing `desc` key produced no error and scored 0.31 against 1.63 for the same item — a 5.3x under-estimate that the end-to-end run reported as success diff --git a/wiki/backend/common/integrations/estimate-derived-thresholds.md b/wiki/backend/common/integrations/estimate-derived-thresholds.md new file mode 100644 index 0000000..8ecc281 --- /dev/null +++ b/wiki/backend/common/integrations/estimate-derived-thresholds.md @@ -0,0 +1,65 @@ +--- +id: backend-common-integrations-estimate-derived-thresholds +domain: backend +category: integrations +applies_to: [general] +confidence: verified +sources: + - https://www.freqtrade.io/en/stable/stoploss/ + - https://en.wikipedia.org/wiki/Slippage_(finance) + - "Live incident 2026-08-13/14 (KIS auto-trading bot): 6 positions exited TAKE_PROFIT at +0.06%–+0.42% against a +2% design; re-anchor in mark_filled() restored the band, regression tests red on pre-fix code" +last_verified: 2026-08-14 +related: [] +--- + +# Absolute Thresholds Derived from a Pre-Execution Estimate + +## When this applies + +You submit an action to an external system whose actual outcome can differ +from the estimate you decided on (market order → fill price vs the quote at +decision time), and you persist **absolute** trigger values derived from that +estimate — stop-loss/take-profit prices, alert thresholds, budget cutoffs. +The actual outcome arrives later as a separate confirmation event (fill +report, webhook, reconciliation). + +## Do this + +1. **Persist the intent as ratios/offsets relative to the anchor** (e.g. + SL −2% / TP +2% of entry), alongside any precomputed absolutes. The ratio + is the durable design value; the absolute is a cache of it. +2. **Re-anchor at the confirmation-recording point.** In the single function + that records the action as confirmed (`mark_filled()`, the fill-webhook + handler), recompute the absolute triggers from the actual outcome, + preserving the stored ratios. Fix it there — not per entry path: entry + paths multiply (daily run, intraday redeploy, manual), while every one of + them funnels through the confirmation recorder. +3. **Treat the estimate/actual gap as expected behavior, not an error.** + Slippage — execution at a price different from the one at decision time — + is a normal property of market orders and widens at opens and in volatile + periods (Wikipedia: Slippage). Reference practice: freqtrade defines + stoploss as a ratio of the entry ("a stoploss of -10% is placed exactly + 10% below the entry point") and places the exchange stoploss order only + after the buy order fills — the anchor follows the fill, not the quote. + +## Edge cases + +| Case | Then | +|------|------| +| Partial fills | Re-anchor on the volume-weighted average fill price — either at each fill event or once on completion; pick one and record which in the recorder | +| The confirmation event can arrive twice (webhook redelivery, reconciliation re-run) | Make the re-anchor idempotent: derive from stored ratios + fill price, never by mutating the previous absolutes incrementally | +| The estimate-derived absolute was already shown or notified to users | Recompute the display from the same anchor (or emit a correction); a UI still quoting the stale absolute contradicts the triggers actually armed | +| The confirmation payload lacks the actual value (no fill price reported) | Keep the estimate-derived values and log loudly that triggers are estimate-anchored; silently treating the estimate as the actual hides the gap this page exists to close | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Fix a stale-anchor bug in each entry path that computes SL/TP | Re-anchor once where the fill is recorded | Entry paths keep multiplying; the confirmation recorder is the funnel they all pass through | +| Keep absolute triggers computed from the pre-order quote after the fill confirms | Recompute from the fill price, preserving the design ratios | Slippage/gap eats the margin: a +2% designed take-profit band collapsed to +0.06%–+0.42% observed triggers, exiting positions for less than fees | + +## Sources + +- https://www.freqtrade.io/en/stable/stoploss/ — stoploss defined as a ratio of the entry price ("a stoploss of -10% is placed exactly 10% below the entry point"); with stoploss-on-exchange, the stoploss order "is placed on the exchange immediately after buy order fills" +- https://en.wikipedia.org/wiki/Slippage_(finance) — execution price differing from the decision-time price is inherent to market orders, larger under volatility +- Live incident 2026-08-13/14, Korea Investment & Securities auto-trading bot: SL/TP absolutes computed from the pre-order signal price survived fills at higher prices; 6 real positions triggered TAKE_PROFIT at +0.06%–+0.42% against a +2% design. Re-anchoring in `mark_filled()` (ratio-preserving recompute from the actual fill price) restored the band; regression tests fail on the pre-fix code and pass after (464-test suite green) diff --git a/wiki/backend/common/integrations/externally-owned-defaults.md b/wiki/backend/common/integrations/externally-owned-defaults.md index 6ab1abf..8617871 100644 --- a/wiki/backend/common/integrations/externally-owned-defaults.md +++ b/wiki/backend/common/integrations/externally-owned-defaults.md @@ -9,7 +9,7 @@ sources: - https://developers.openai.com/api/docs/api-reference/models/list - https://docs.litellm.ai/docs/proxy/model_discovery last_verified: 2026-07-31 -related: [backend-common-llm-completion-response-validation, infrastructure-config-environment-config, qa-process-release-gates] +related: [backend-common-llm-completion-response-validation, infrastructure-config-environment-config, qa-process-release-gates, backend-common-integrations-consumer-required-fields] --- # Defaults That Name a Resource Owned Outside the Repository diff --git a/wiki/backend/common/integrations/robots-txt-and-source-selection.md b/wiki/backend/common/integrations/robots-txt-and-source-selection.md index 70086d7..ff735cf 100644 --- a/wiki/backend/common/integrations/robots-txt-and-source-selection.md +++ b/wiki/backend/common/integrations/robots-txt-and-source-selection.md @@ -11,7 +11,7 @@ sources: - https://housing.seoul.go.kr/robots.txt - https://apply.gh.or.kr/robots.txt last_verified: 2026-08-05 -related: [backend-common-integrations-externally-owned-defaults, backend-common-reliability-timeouts-and-retries] +related: [backend-common-integrations-externally-owned-defaults, backend-common-reliability-timeouts-and-retries, qa-environments-headless-browser-bot-blocking] --- # Choosing a Source to Crawl by Reading robots.txt diff --git a/wiki/backend/common/ml/mape-aligned-point-prediction.md b/wiki/backend/common/ml/mape-aligned-point-prediction.md new file mode 100644 index 0000000..ce45ba4 --- /dev/null +++ b/wiki/backend/common/ml/mape-aligned-point-prediction.md @@ -0,0 +1,61 @@ +--- +id: backend-common-ml-mape-aligned-point-prediction +domain: backend +category: ml +applies_to: [general] +confidence: verified +sources: + - https://arxiv.org/abs/0912.0902 + - https://arxiv.org/abs/1605.02541 +last_verified: 2026-08-07 +related: [] +--- + +# Point Predictions from a Median-Predicting Model Scored by MAPE + +## When this applies + +A regression model is evaluated by MAPE (mean absolute percentage error) and was +trained to predict the conditional median — the standard log-target + L1-loss +setup (LightGBM/CatBoost `regression_l1` on `log(y)`, quantile q50, or any +median regression). The model looks well-fitted yet systematically overshoots +on MAPE, or you are deciding what point value to emit from such a model. + +## Do this + +The MAPE-optimal point prediction is not the conditional median. MAPE is +1/y-weighted absolute error, so its Bayes rule is the median of the predictive +distribution reweighted by 1/y (Gneiting's med^(-1)); for a lognormal +conditional distribution that equals `median × exp(−σ²)`. A median-predicting +model is therefore a *structural* overpredictor under MAPE, and the gap grows +with conditional variance. + +| Case | Do | +|------|----| +| Emitting point predictions under a MAPE metric | Apply a per-row correction `pred × exp(−λσ²)` instead of shipping the raw median | +| Estimating per-row σ | Train quantile models (e.g. q16/q84) and use half the spread in log space: `σ ≈ (log(q84) − log(q16)) / 2` | +| Choosing λ | Start from λ = 0.5, not the theoretical 1.0, and select by holdout consistency across several evaluation periods — accept a λ only if it improves every period, not just the average | +| Deciding between global and per-row correction | Per-row: a single global constant assumes homoscedasticity and undercorrects high-variance rows while overcorrecting low-variance ones | + +## Edge cases + +| Case | Then | +|------|------| +| The theoretical λ = 1.0 degrades some holdout periods while improving others | Expected — quantile-spread σ estimates are inflated (quantile crossings, spread widened by estimation noise), so full-strength shrinkage overshoots on low-variance segments; back λ off until every period improves | +| The conditional distribution is far from lognormal (heavy point mass, multimodal) | The `exp(−σ²)` form no longer follows; fall back to selecting a multiplicative shrinkage factor purely by holdout search | +| The metric is MAE or RMSE, not MAPE | Do not shrink — the median (MAE) or mean (RMSE) is already the optimal point forecast; this correction only applies to relative-error metrics | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Ship raw log-target L1 predictions because the model "predicts the center" | Apply the `exp(−λσ²)` shrinkage before scoring | Under MAPE the center that scores best is the 1/y-weighted median, which sits below the ordinary median | +| Apply the textbook λ = 1.0 because the derivation says so | Validate λ ∈ [0.5, 1.0] on multi-period holdout and keep the value that improves all periods | The derivation assumes σ is exact; a spread-estimated σ is biased upward, so the theory value overcorrects | +| Fix overprediction with one global scale factor tuned on the mean | Use the per-row σ from quantile models | Global scaling ignores heteroscedasticity — precisely the rows where the median-vs-MAPE gap is largest get the wrong correction | + +## Sources + +- https://arxiv.org/abs/0912.0902 — Gneiting, "Making and Evaluating Point Forecasts" (JASA 106:746–762, 2011): Table 5 gives the Bayes rule per scoring function; for absolute percentage error it is the β-median with β = −1, the median of the 1/y-reweighted predictive distribution +- https://arxiv.org/abs/1605.02541 — de Myttenaere et al., "Mean Absolute Percentage Error for regression models" (Neurocomputing 2016): MAPE-optimal regression is equivalent to 1/y-weighted MAE regression +- Lognormal algebra: for Y ~ LN(μ, σ²), the 1/y-reweighted density is LN(μ − σ², σ²), whose median is `exp(μ − σ²)` = `median(Y) × exp(−σ²)` +- Field validation 2026-08-07 (Seoul commercial-building AVM, LightGBM + CatBoost, log target + L1): λ = 0.5 with q16/q84-spread σ improved MAPE on all 4 holdout years (19.18% → 18.86% excluding 2025 outliers); λ = 1.0 degraded 2 of 4 years — the λ selection practice is field-tested, the shrinkage mechanism itself is the sourced theory above diff --git a/wiki/backend/common/orm/transaction-boundaries.md b/wiki/backend/common/orm/transaction-boundaries.md index d4d067f..c3dc57d 100644 --- a/wiki/backend/common/orm/transaction-boundaries.md +++ b/wiki/backend/common/orm/transaction-boundaries.md @@ -9,7 +9,7 @@ sources: - https://docs.spring.io/spring-framework/reference/data-access/transaction/declarative/tx-propagation.html - https://vladmihalcea.com/spring-transaction-best-practices/ last_verified: 2026-07-10 -related: [backend-common-jobs-idempotent-handlers, backend-common-errors-exception-handling, databases-transactions-isolation-level-selection] +related: [backend-java-jpa-raw-jdbc-inside-a-jpa-transaction, backend-common-jobs-idempotent-handlers, backend-common-errors-exception-handling, databases-transactions-isolation-level-selection] --- # Transaction Boundaries in Application Code diff --git a/wiki/backend/common/reliability/timeouts-and-retries.md b/wiki/backend/common/reliability/timeouts-and-retries.md index f176013..e15dc46 100644 --- a/wiki/backend/common/reliability/timeouts-and-retries.md +++ b/wiki/backend/common/reliability/timeouts-and-retries.md @@ -9,7 +9,7 @@ sources: - https://sre.google/sre-book/addressing-cascading-failures/ - https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/ last_verified: 2026-08-05 -related: [backend-common-api-design-idempotency, backend-common-llm-completion-response-validation, backend-common-reliability-client-side-rate-limiting, debugging-concurrency-intermittent-failures] +related: [backend-java-jpa-raw-jdbc-inside-a-jpa-transaction, backend-common-api-design-idempotency, backend-common-llm-completion-response-validation, backend-common-reliability-client-side-rate-limiting, debugging-concurrency-intermittent-failures] --- # Calling Another Service over the Network: Timeouts, Retries, Backoff diff --git a/wiki/backend/index.md b/wiki/backend/index.md index fa04bae..81e97e1 100644 --- a/wiki/backend/index.md +++ b/wiki/backend/index.md @@ -5,10 +5,10 @@ three stack subtrees — route by concern first, stack second: | Subtree | Route there when | |---------|------------------| -| [common](#common-language-agnostic) (below) | The concern is language-agnostic: API contracts, enumerating call sites before a contract change, idempotency, JWT issuance, outbound calls, caching, jobs, transactions in app code, shared state/pools, exception structure, consuming LLM APIs (completion validation, context budgeting), consuming external-API responses, externally-owned defaults, object-storage references, sync-vs-async integration choice, WebSocket/SSE connection lifecycle | +| [common](#common-language-agnostic) (below) | The concern is language-agnostic: API contracts, enumerating call sites before a contract change, idempotency, JWT issuance, outbound calls, caching, jobs, transactions in app code, shared state/pools, exception structure, consuming LLM APIs (completion validation, context budgeting), MAPE-aligned point-prediction calibration, consuming external-API responses, externally-owned defaults, object-storage references, sync-vs-async integration choice, WebSocket/SSE connection lifecycle | | [java](java/index.md) | You are writing/reviewing JVM backend code (Java/Kotlin, Spring, JPA/Hibernate) and the concern is stack-specific: entity mapping, persistence context, proxy pitfalls, JVM threads/memory | | [node](node/index.md) | You are writing/reviewing Node.js/TypeScript backend code: event-loop blocking, promise error handling, runtime validation at boundaries, graceful shutdown | -| [python](python/index.md) | You are writing/reviewing Python backend code: GIL/concurrency model, pydantic validation, WSGI/ASGI workers, language traps | +| [python](python/index.md) | You are writing/reviewing Python backend code: GIL/concurrency model, pydantic validation, WSGI/ASGI workers, language traps, packaging data files and resolving them after install | Load a stack page IN ADDITION to the matching common page when both apply — common owns the principle, the stack page owns the mechanics. SQL, index, and DB-side @@ -33,7 +33,12 @@ Match your situation to a "load when" line; load only matching pages. | Page | Load when | |------|-----------| +| [widening-a-closed-value-table](common/change-impact/widening-a-closed-value-table.md) | Adding an entry to a closed table mapping names to magnitudes or codes (duration units, status codes, currency exponents, severity levels) that lives as a named constant; scoping that change from a search for the constant's name; a new entry parses at one layer and is rejected or mis-converted at another; deciding what to do about an inlined copy of the table in a hot path, a second language backend, or a fixture | +| [corpus-sweep-before-a-rejection-rule](common/change-impact/corpus-sweep-before-a-rejection-rule.md) | Adding a rule to a compiler/linter/parser/schema validator/repo gate that will start rejecting input the tool accepted silently, and the existing corpus must keep passing; producing the evidence a plan needs before writing the rule (reject count + rejected-path list, enumeration method stated); such a rule landed and went red on inputs nobody had called defective; deciding between narrowing the rule, an opt-in strictness level, and an exemption (whether unimplemented declarative input should reject/warn/ignore at all → common/api-design/unenforced-declarations) | +| [aggregation-layer-of-a-shared-helper](common/change-impact/aggregation-layer-of-a-shared-helper.md) | A plan, brief, or review comment says to unify or replace "the N call sites" of a helper and names them by line number rather than by what the code does; one site feeds a set-level SQL aggregate while another returns a per-row value that application code reduces later; a "shared helper" landed with one consumer adopting it and the other keeping its old semantics; deciding which layer owns a missing-value rule and whether the plan's grep-count acceptance criterion is reachable at all | | [call-site-enumeration](common/change-impact/call-site-enumeration.md) | Changing the contract of a function/method/constructor other code calls — adding, removing, reordering or redefining a parameter — and you need the complete call-site list; scoping such a migration from a search; a migration scoped from recon came back green and then failed on call sites the search never listed; deciding whether to append a parameter or make it keyword-only (release-level re-test scope → qa/process/regression-scope) | +| [cross-module-consumer-census](common/change-impact/cross-module-consumer-census.md) | Your task added a public function, endpoint, export, or hook whose consumer belongs to a *different* task (parallel work split by file ownership, a backend change whose UI wiring is another ticket); deciding whether that task is done; a feature typed, tested, reviewed and merged changes nothing at runtime; separating same-module helpers and entry points from genuine orphans in a zero-consumer list | +| [inserting-a-guard-before-an-existing-side-effect](common/change-impact/inserting-a-guard-before-an-existing-side-effect.md) | Implementing, adopting, or auditing a planned change that adds a precondition guard to an existing script/function where the plan names the insertion point in prose ("after X is built"); the target has an earlier unconditional default/auto-create side effect touching the state the guard checks; writing the test for such a guard's refusal path | ### reliability @@ -60,6 +65,7 @@ Match your situation to a "load when" line; load only matching pages. | Page | Load when | |------|-----------| | [exception-handling](common/errors/exception-handling.md) | Writing a catch block or deciding where errors are handled/logged/translated in a service — catch placement, log-once, wrapping with cause preserved, typed results for expected outcomes; one fault producing duplicate alerts | +| [diagnostics-from-a-shared-code-path](common/errors/diagnostics-from-a-shared-code-path.md) | Writing or reviewing a rejection/validation message emitted by one function several callers reach (a check shared by two syntaxes, request and response, two config blocks) — especially when the message names a construct in literal text or tells the author what to write instead; a user reports a rejection naming a construct they did not write | | [async-failure-handling](common/errors/async-failure-handling.md) | Handing work to in-process async (@Async, unawaited futures/promises) — deciding fire-and-forget vs consumed future vs durable job; side effects silently never happening with no error logs; unobserved futures; async work enqueued inside a transaction | ### auth @@ -88,12 +94,20 @@ Match your situation to a "load when" line; load only matching pages. | [completion-response-validation](common/llm/completion-response-validation.md) | Consuming OpenAI-compatible `/chat/completions` output as a final artifact (summary, document, notification); LLM responses coming back empty or truncated while HTTP status is 200; a reasoning-family model may be routed onto the alias you call | | [context-window-budget](common/llm/context-window-budget.md) | Repointing an LLM client or agent CLI at a different model, a self-hosted server (vLLM/Ollama), or a gateway (LiteLLM); setting `max_tokens` for a client whose default was sized for a larger model; the first request after such a switch returns 400 with a context-window error; deciding where to set the cap (request body vs client env var vs gateway config) and how to point the base URL at a proxy; handling truncation that arrives as a normal 200 | +### ml + +| Page | Load when | +|------|-----------| +| [mape-aligned-point-prediction](common/ml/mape-aligned-point-prediction.md) | A regression model evaluated by MAPE was trained as a median predictor (log target + L1 loss, or quantile q50) — deciding what point value to emit, or the model systematically overpredicts on MAPE despite fitting well; choosing between a global scale factor and per-row variance-based correction | + ### integrations | Page | Load when | |------|-----------| | [externally-owned-defaults](common/integrations/externally-owned-defaults.md) | A code/config default names a resource the repo does not own (model alias, endpoint, bucket, queue, index) — reviewing or merging a PR that claims that default works, adding a startup check that the name still resolves, or diagnosing a default path that broke with no code change | +| [consumer-required-fields](common/integrations/consumer-required-fields.md) | Writing an adapter that maps one module's records into the payload a second module (scoring engine, plugin, external client) consumes, with the target shape taken from a docstring, README example, or sample payload; such an adapter runs end to end with no error and the downstream numbers come out low; deciding which mapped fields need their own assertion | | [robots-txt-and-source-selection](common/integrations/robots-txt-and-source-selection.md) | Choosing which site to fetch a published dataset from and reading its robots.txt to decide whether your client may crawl it; the file contains a `Disallow: /` and you are deciding whose group it belongs to; setting the crawler's User-Agent and checking that token against the file; robots.txt returned a non-200 status; the origin restricts your token and you are looking for a portal that republishes the same records | +| [estimate-derived-thresholds](common/integrations/estimate-derived-thresholds.md) | Submitting an action to an external system whose actual outcome can differ from the decision-time estimate (market order fill vs quote) while persisting absolute trigger values derived from that estimate (SL/TP prices, alert thresholds); derived triggers fire immediately or at the wrong level right after the action confirms; choosing where to recompute them from the actual outcome | ### storage diff --git a/wiki/backend/java/index.md b/wiki/backend/java/index.md index 8ac2e14..3247979 100644 --- a/wiki/backend/java/index.md +++ b/wiki/backend/java/index.md @@ -15,7 +15,9 @@ load it alongside the stack page here — these pages link the exact ids. | Page | Load when | |------|-----------| | [entity-mapping](jpa/entity-mapping.md) | Writing or reviewing JPA entity classes/associations — fetch types (to-one EAGER default), bidirectional sync helpers, entity equals/hashCode; debugging `LazyInitializationException`, `MultipleBagFetchException`, entities vanishing from Sets, or unexpected joins/queries traced to mappings; deciding DTO projection vs entity for read-only endpoints; evaluating Open Session in View | +| [raw-jdbc-inside-a-jpa-transaction](jpa/raw-jdbc-inside-a-jpa-transaction.md) | One query in a JPA/Hibernate service was dropped to `JdbcTemplate`/`NamedParameterJdbcTemplate` and `@Transactional(timeout = N)` is what bounds it; one endpoint's slow queries are cancelled at N seconds on some paths and run for minutes on others; the same timeout surfaces as a 4xx on one path and a 5xx on another; writing the handler branch for a query-cancellation exception, or deciding where to put a statement timeout | | [persistence-context](jpa/persistence-context.md) | Debugging changes saved without calling save (dirty checking), stale reads within one transaction (first-level cache), flush timing surprises around queries, detached-entity errors (merge vs persist, lost updates after merge); designing or fixing slow/memory-hungry JPA batch inserts; choosing IDENTITY vs SEQUENCE id generation for batch-heavy tables | +| [not-null-check-and-lifecycle-callbacks](jpa/not-null-check-and-lifecycle-callbacks.md) | A write fails with `PropertyValueException: not-null property references a null or transient value` and you must attribute it to an attribute/path (dotted path vs plain name, embeddables, nulled transient references); about to fix such a failure in a `@PreUpdate` listener; asking why an entity's `nullable = false` is (or stopped being) enforced at runtime — `hibernate.check_nullability` vs Bean Validation on the classpath; deciding between `@Column(nullable = false)`, `@NotNull`, and the DB constraint as the enforcing layer | ## spring diff --git a/wiki/backend/java/jpa/not-null-check-and-lifecycle-callbacks.md b/wiki/backend/java/jpa/not-null-check-and-lifecycle-callbacks.md new file mode 100644 index 0000000..2b2640a --- /dev/null +++ b/wiki/backend/java/jpa/not-null-check-and-lifecycle-callbacks.md @@ -0,0 +1,109 @@ +--- +id: backend-java-jpa-not-null-check-and-lifecycle-callbacks +domain: backend +category: jpa +applies_to: [java, kotlin, jpa, hibernate, spring] +confidence: verified +sources: + - https://github.com/hibernate/hibernate-orm/blob/5.6/hibernate-core/src/main/java/org/hibernate/engine/internal/Nullability.java + - https://github.com/hibernate/hibernate-orm/blob/6.6/hibernate-core/src/main/java/org/hibernate/event/internal/DefaultFlushEntityEventListener.java + - https://github.com/hibernate/hibernate-orm/blob/6.6/hibernate-core/src/main/java/org/hibernate/action/internal/AbstractEntityInsertAction.java + - https://github.com/hibernate/hibernate-orm/blob/6.6/hibernate-core/src/main/java/org/hibernate/boot/beanvalidation/TypeSafeActivator.java + - https://docs.hibernate.org/orm/6.6/javadocs/org/hibernate/cfg/ValidationSettings.html + - https://docs.hibernate.org/orm/6.6/javadocs/org/hibernate/boot/spi/SessionFactoryOptions.html + - https://thorben-janssen.com/hibernate-tips-whats-the-difference-between-column-nullable-false-and-notnull/ + - https://www.baeldung.com/hibernate-not-null-error +last_verified: 2026-08-13 +related: [backend-java-jpa-persistence-context, backend-java-kotlin-frameworks-and-jpa, databases-schema-design-nullability-and-defaults, databases-data-survey-audit-columns-as-update-evidence] +--- + +# Attributing Hibernate's Not-Null Check Exception, and Knowing Whether It Runs + +## When this applies + +A write fails with `PropertyValueException: not-null property references a null or +transient value : ` and you must decide which attribute and which code path +produced it; you are about to fix it inside a `@PreUpdate` listener; or you are asking +why an entity's `nullable = false` was never enforced before (or stopped being). + +Column-side nullability design → [databases-schema-design-nullability-and-defaults]. +Reading the failed rows' audit columns → [databases-data-survey-audit-columns-as-update-evidence]. + +## Do this + +1. Identify the attribute from the **path shape**, not from the words "or transient". + `Nullability` throws this message from two sites that share one hardcoded literal — + it is the only production occurrence of that string in hibernate-orm — so the + wording says nothing about which site fired: + +| Path in the message | Means | Inspect | +|---------------------|-------|---------| +| No dot (`title`) | A top-level attribute of the entity held `null` when the check ran | That attribute's column mapping, and every path that assembles the entity | +| Contains a dot (`address.city`) | A sub-attribute of a composite value (`@Embeddable`) held `null`; only `checkSubElementsNullability` → `buildPropertyPath` produces dots, and it recurses into `CompositeType` | The embeddable's own `nullable = false` attributes — the owning entity attribute (`address`) was non-null | +| Contains a dot and the parent attribute is a collection | The same, reached through a collection whose **element type** is composite (`@ElementCollection` of `@Embeddable`); the first loaded non-null element decides | The element class's not-null attributes | + +2. On the UPDATE path, drop `@PreUpdate`/`@PostUpdate` from both the suspect list and + the fix. `DefaultFlushEntityEventListener.scheduleUpdate` runs + `new Nullability( session ).checkNullability( values, persister, … )` and only then + adds `EntityUpdateAction` to the action queue; the callbacks fire inside that + action's `execute()`. The order is the same in 5.6, 6.6 and 7.0. So when this + exception is thrown, no `@PreUpdate` listener has run on that entity: a listener + cannot have nulled the value, and a listener cannot supply it. +3. Fix the value where the entity state is assembled — the service, mapper, or + deserializer that produced the instance — or change the declared nullability if + "absent" is a real state of the domain. +4. Before ruling an attribute out, check whether the loop even examined it. + `Nullability` skips an attribute when it is not insertable (INSERT) or not + updatable (UPDATE), when its value is `UNFETCHED_PROPERTY` (lazy, not loaded), and + when Hibernate generates the value in memory (`GenerationTiming != NEVER` — e.g. + `@CreationTimestamp`, `@UpdateTimestamp`). +5. Measure whether the check is active instead of inferring it from the mapping. + `hibernate.check_nullability` "Defaults to disabled if Bean Validation is present in + the classpath and annotations are used, or enabled otherwise": + `TypeSafeActivator.applyCallbackListeners` calls `setCheckNullability( false )` + whenever the validation mode is `CALLBACK`/`AUTO` **and** the setting has no value. + Adding or removing a dependency such as `spring-boot-starter-validation` therefore + flips it. Read it, or assert it: + +| To establish | Do | +|--------------|----| +| The effective setting at startup | `emf.unwrap( SessionFactoryImplementor.class ).getSessionFactoryOptions().isCheckNullability()` | +| That the behaviour holds for this build | A test that flushes an entity whose `nullable = false` attribute is null and expects `PropertyValueException` | + +6. When the app-level check must hold regardless of which dependencies are on the + classpath, set `hibernate.check_nullability=true` explicitly. `@Column(nullable = + false)` alone does not give you a Bean Validation constraint — it "adds a not null + constraint to the database column, if Hibernate creates the database table + definition" — so with the core check off, the enforcement left is the DB constraint. + Add `@NotNull` (Kotlin: `@field:NotNull` → + [backend-java-kotlin-frameworks-and-jpa]) when you want the validator to reject it. + +## Edge cases + +| Case | Then | +|------|------| +| The same message on `persist()`/INSERT | The check runs in `AbstractEntityInsertAction.nullifyTransientReferencesIfNotAlready()`, immediately after `nullifyTransientReferences( getState() )` — a to-one attribute holding an unsaved instance is nulled first, then reported as null. That is where "or transient" comes from; the reported path is still the plain attribute name | +| The path names an attribute whose DB column is nullable | The entity declares not-null while the column allows NULL; existing rows can already violate it, and they surface only once this check is on (directive 5) | +| No exception at all despite a null on a `nullable = false` attribute | The check is off — the statement reaches the database, where a real NOT NULL constraint raises a `ConstraintViolationException` and a nullable column accepts the row silently | +| The failing attribute is `@UpdateTimestamp`/`@CreationTimestamp` | It is skipped by the check (directive 4); the reported path belongs to another attribute | +| The write path is a bulk JPQL/HQL `update` or native SQL | No entity flush happens, so this check never runs — the database constraint is the only gate → [databases-data-survey-audit-columns-as-update-evidence] | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Add a `@PreUpdate` listener that fills the missing value | Set it where the entity state is assembled (directive 3) | The check runs before `EntityUpdateAction` exists, so no `@PreUpdate` has run — the listener is never reached on the failing flush | +| Read "or transient" as evidence that an unsaved association caused it | Read the path shape (directive 1); on INSERT, treat a nulled transient reference as one of the causes | One literal is shared by both throw sites; the distinct unsaved-instance error carries its own message, "object references an unsaved transient instance" | +| Conclude the check is on because the mapping says `nullable = false` | Read `isCheckNullability()` or assert the exception in a test (directive 5) | Bean Validation on the classpath disables the core check unless the setting is explicit | +| Set `hibernate.check_nullability=false` to get past the exception | Supply the value, or relax the declared nullability | Disabling moves the failure to the DB constraint, or writes the incomplete row when the column is nullable | + +## Sources + +- https://github.com/hibernate/hibernate-orm/blob/5.6/hibernate-core/src/main/java/org/hibernate/engine/internal/Nullability.java — two throw sites share the literal `"not-null property references a null or transient value"`; the dotted path comes only from `buildPropertyPath(...)` via `checkSubElementsNullability`, which recurses into `CompositeType` and into collections whose element type is composite; the loop skips non-checkable, `UNFETCHED_PROPERTY`, and `GenerationTiming != NEVER` attributes; comment: "Typically when Bean Validation is on, we don't want to validate null values at the Hibernate Core level. Hence the checkNullability setting." +- https://github.com/hibernate/hibernate-orm/blob/6.6/hibernate-core/src/main/java/org/hibernate/event/internal/DefaultFlushEntityEventListener.java — `scheduleUpdate`: "check nullability but do not doAfterTransactionCompletion command execute" → `new Nullability( session ).checkNullability(...)` precedes `new EntityUpdateAction(...)` (same order on 5.6 and 7.0) +- https://github.com/hibernate/hibernate-orm/blob/6.6/hibernate-core/src/main/java/org/hibernate/action/internal/AbstractEntityInsertAction.java — `nullifyTransientReferencesIfNotAlready()` nullifies transient references and then runs the CREATE-type nullability check +- https://github.com/hibernate/hibernate-orm/blob/6.6/hibernate-core/src/main/java/org/hibernate/boot/beanvalidation/TypeSafeActivator.java — "de-activate not-null tracking at the core level when Bean Validation is present unless the user explicitly asks for it": guarded by validation mode `CALLBACK`/`AUTO`, then `if ( cfgService.getSettings().get( CHECK_NULLABILITY ) == null ) … setCheckNullability( false )` +- https://docs.hibernate.org/orm/6.6/javadocs/org/hibernate/cfg/ValidationSettings.html — `CHECK_NULLABILITY`: "Enable nullability checking, raises an exception if an attribute marked as not null is null at runtime"; "Defaults to disabled if Bean Validation is present in the classpath and annotations are used, or enabled otherwise" +- https://docs.hibernate.org/orm/6.6/javadocs/org/hibernate/boot/spi/SessionFactoryOptions.html — `boolean isCheckNullability()` exposes the effective setting +- https://thorben-janssen.com/hibernate-tips-whats-the-difference-between-column-nullable-false-and-notnull/ — `@Column(nullable = false)` "adds a not null constraint to the database column, if Hibernate creates the database table definition" and otherwise leaves validation to the database; `@NotNull` is what Bean Validation checks on pre-persist/pre-update +- https://www.baeldung.com/hibernate-not-null-error — the widely repeated two-cause framing this page corrects: the same message attributed to "a null value for a column marked with nullable = false" and to "an association referencing an unsaved instance", with no mention of the path shape diff --git a/wiki/backend/java/jpa/raw-jdbc-inside-a-jpa-transaction.md b/wiki/backend/java/jpa/raw-jdbc-inside-a-jpa-transaction.md new file mode 100644 index 0000000..b912f59 --- /dev/null +++ b/wiki/backend/java/jpa/raw-jdbc-inside-a-jpa-transaction.md @@ -0,0 +1,124 @@ +--- +id: backend-java-jpa-raw-jdbc-inside-a-jpa-transaction +domain: backend +category: jpa +applies_to: [java, spring, jpa, hibernate] +confidence: verified +sources: + - https://docs.spring.io/spring-framework/docs/current/javadoc-api/org/springframework/orm/jpa/JpaTransactionManager.html + - https://docs.spring.io/spring-framework/docs/current/javadoc-api/org/springframework/jdbc/datasource/DataSourceUtils.html + - https://github.com/spring-projects/spring-framework/blob/v6.2.0/spring-orm/src/main/java/org/springframework/orm/jpa/JpaTransactionManager.java + - https://github.com/spring-projects/spring-framework/blob/v6.2.0/spring-jdbc/src/main/java/org/springframework/jdbc/support/SQLStateSQLExceptionTranslator.java + - https://github.com/hibernate/hibernate-orm/blob/main/hibernate-core/src/main/java/org/hibernate/dialect/PostgreSQLDialect.java + - https://mybatis.org/mybatis-3/configuration.html + - https://github.com/mybatis/mybatis-3/blob/273ec6508c348513456cb2cc48f943b28196a56d/src/main/java/org/apache/ibatis/executor/statement/BaseStatementHandler.java + - https://github.com/mybatis/mybatis-3/blob/273ec6508c348513456cb2cc48f943b28196a56d/src/main/java/org/apache/ibatis/executor/statement/StatementUtil.java + - https://github.com/postgres/postgres/blob/083ac033419f690758508e08c1736089384bbee8/src/backend/tcop/postgres.c +last_verified: 2026-08-11 +related: + [ + backend-common-orm-transaction-boundaries, + backend-common-reliability-timeouts-and-retries, + backend-common-errors-exception-handling, + backend-java-spring-proxy-pitfalls, + ] +--- + +# A Raw JdbcTemplate Query Inside a JPA-Managed Transaction + +## When this applies + +A Spring service is JPA/Hibernate-backed, and one query was dropped to +`JdbcTemplate`/`NamedParameterJdbcTemplate` for performance while +`@Transactional(timeout = N)` is what is supposed to bound its runtime. Also when +one endpoint's slow queries are cancelled at N seconds through some paths and run +for minutes through others, or when the same timeout surfaces as a 4xx on one +path and a 5xx on another. + +Where the transaction boundary belongs → [backend-common-orm-transaction-boundaries]. +The annotation having no effect at all (self-invocation, non-public method) → +[backend-java-spring-proxy-pitfalls]. + +## Do this + +1. **Measure the raw path's actual cancellation point before trusting the + declared timeout.** Run a query you know exceeds N through that exact method + and record the elapsed time (a JDBC-level logger such as p6spy, or the + database's own cancellation message). The declared timeout and the applied + timeout are separate facts, and nothing logs the gap between them. + +2. **Trace the deadline's path.** `JdbcTemplate` ends every statement setup with + `DataSourceUtils.applyTimeout(stmt, getDataSource(), getQueryTimeout())`, which + applies "the current transaction timeout, **if any**". It looks the deadline up + as a `ConnectionHolder` bound to *its own* `DataSource` instance; with no + holder it falls back to the template's own `queryTimeout`, which defaults to + `-1` and so sets nothing. Check each link: + +| Link | Check | If it fails | +|------|-------|-------------| +| The transaction manager knows the DataSource | `JpaTransactionManager` binds a `ConnectionHolder` only inside `if (getDataSource() != null)`; it "will autodetect the DataSource used as the connection factory of the EntityManagerFactory" | Set it explicitly (`setDataSource`), matching the EntityManagerFactory's DataSource | +| The JpaDialect can expose the JDBC connection | The bind is skipped when `getJpaDialect().getJdbcConnection(...)` returns `null` — `DefaultJpaDialect`'s implementation returns `null`; enable debug logging on the manager and look for "Not exposing JPA transaction … because JpaDialect … does not support JDBC Connection retrieval" | Configure the vendor dialect (`HibernateJpaDialect`), which the docs state "requires a vendor-specific `JpaDialect` to be configured" | +| The template uses the same DataSource instance | The lookup key is the `DataSource` object the template holds; a second `DataSource` bean, or one wrapped after the manager captured it, is a different key | Build the template from the same bean, or from a `TransactionAwareDataSourceProxy` | +| The timeout is declared where the proxy sees it | `timeout` on the annotation only reaches `doBegin` when that call is proxied | → [backend-java-spring-proxy-pitfalls] | + +3. **Give the raw path a deadline that does not depend on that chain.** Set + `setQueryTimeout(N)` on the template, or set the database's own statement + timeout for the connection, so the bound is present whether or not the holder + is. `applyTimeout` prefers the transaction's remaining time when a holder + exists, so an explicit value is a floor, not a conflict. + +4. **Handle both timeout exception types before shipping the fix.** Cancellation + raises different Spring exceptions per path, so a handler branch written for + the JPA path does not cover the raw one: + +| Path | Chain | Spring exception | +|------|-------|------------------| +| Hibernate/JPA query | PostgreSQL SQLState `57014` → `org.hibernate.QueryTimeoutException` → `HibernateJpaDialect` | `org.springframework.dao.QueryTimeoutException` | +| Raw JdbcTemplate, Spring Framework ≤ 6.2.x | `PSQLException` (a plain `SQLException`, so the JDBC-4 `SQLTimeoutException` branch does not match) → SQLState class `57` | `DataAccessResourceFailureException` | +| Raw JdbcTemplate, Spring Framework ≥ 7.0.0 | Same chain, plus a `"57014".equals(sqlState)` check ahead of the class-57 mapping | `org.springframework.dao.QueryTimeoutException` | + +5. **Keep the timeout's response status the same on both paths, and fix the + missing bound as its own change.** A path that now runs for minutes and a path + cancelled at N seconds are one defect wearing two status codes; the status + difference is the symptom that leads to it. + +## Edge cases + +| Case | Then | +|------|------| +| A server-side `statement_timeout` is also set | It bounds every path independently of Spring, which makes it the cheapest guard to add; keep the application timeout below it so the application's own error is what surfaces | +| The method is `@Transactional(readOnly = true)` | The timeout is unaffected by `readOnly` — both are set on the same transaction object — so a read-only annotation is not evidence the deadline applies | +| The raw query runs outside any transaction (no annotation on the path) | There is no holder to carry a deadline at all, so the explicit `setQueryTimeout` from step 3 is the only bound | +| The team's first fix is to map the new exception to a 4xx so the alerts stop | Map it after the bound exists — otherwise minute-long queries leave the 5xx alerting entirely and the remaining defect has no signal | +| The database is MySQL rather than PostgreSQL | The class-57 mapping is PostgreSQL's SQLState; Spring's fallback translator also returns `QueryTimeoutException` when the driver's exception class name contains "Timeout", which is the MySQL path — confirm which branch your driver takes before writing the handler | +| The same service also runs queries through MyBatis | MyBatis carries a bound that does not depend on the holder chain at all: `BaseStatementHandler.setStatementTimeout` takes the mapped statement's `timeout` attribute, else the global `defaultStatementTimeout`, calls `stmt.setQueryTimeout(...)`, and only then narrows to the transaction's remaining time via `StatementUtil.applyTransactionTimeout`. So a MyBatis path can be bounded on the same request where the raw-JdbcTemplate path is not — enumerate the access technologies per path, not per service | +| You have a measured duration but not the layer that cancelled it | Fingerprint it against every timeout value configured in the application. A duration landing just above one configured value identifies that layer (measured: 120,010 ms against a MyBatis `defaultStatementTimeout` of 120; 10,012 ms against `@Transactional(timeout = 10)`). A duration matching no configured value **and** differing run to run (151,558 ms, then 163,489 ms) means no timeout applied on that path | +| The database log reads `canceling statement due to user request`, not `… due to statement timeout` | Read it as "a client-side deadline fired", then fingerprint which one. PostgreSQL emits `user request` as the fall-through branch after excluding lock timeout, statement timeout and autovacuum — i.e. an external cancel such as the JDBC driver acting on `setQueryTimeout`, or `pg_cancel_backend()`. Both messages carry SQLSTATE 57014, so only the message text separates a server-side `statement_timeout` from an application-side bound | +| A `sql-error-codes.xml` file sits at the classpath root | The template then uses `SQLErrorCodeSQLExceptionTranslator` instead of the subclass/state chain above, so re-derive the exception type for your file's mappings | +| The same value is also enforced by an HTTP or gateway timeout | The client-visible failure comes from whichever fires first; order them so the database cancellation happens first, or the query keeps running after the response is gone ([backend-common-reliability-timeouts-and-retries]) | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Read `@Transactional(timeout = 10)` on the class as evidence every query inside is bounded at 10s | Measure one deliberately-slow query per access path | Measured on one endpoint: the Hibernate path cancelled at 10,012 ms while raw-JdbcTemplate calls on the same annotated path ran 151,558 ms and 163,489 ms | +| Treat the new `DataAccessResourceFailureException` as a newly-broken dependency | Read it as the same timeout arriving through the other translator branch | SQLState class `57` is in Spring's `DATA_ACCESS_RESOURCE_FAILURE_CODES` through 6.2.x; the message still carries the cancellation text | +| Add the raw path's exception to the 4xx branch and close the incident | Add the missing query timeout, then align the status | The status mapping removes the alert; the unbounded query is what the alert was pointing at | +| Wrap the raw call in an application-level watchdog (a future with a timeout) | Set the statement timeout so the database cancels the work | A cancelled wrapper returns control while the query keeps running and holds its connection | +| Assume the DataSource is wired because Spring Boot auto-configured the manager | Verify the holder exists — check the manager's debug log line, or assert that a `@Transactional(timeout = 1)` method's raw query fails | Autodetection covers the DataSource, and the bind still fails silently when the dialect cannot expose the connection | +| Conclude the path has no timeout because the statement ran for minutes | Fingerprint the duration against every value configured in the application first | Another access technology's own default may be bounding it at a number you did not look for; only a duration that matches nothing *and* varies run to run rules all of them out | +| Rely on the Spring 7 mapping and write one handler branch for `QueryTimeoutException` | Pin the framework version the branch assumes, and keep the `DataAccessResourceFailureException` branch while any service is on 6.2.x or earlier | The `"57014"` check exists in 7.0.0 and is absent in 6.2.8 and every earlier tag checked | + +## Sources + +- https://docs.spring.io/spring-framework/docs/current/javadoc-api/org/springframework/orm/jpa/JpaTransactionManager.html — "This transaction manager also supports direct DataSource access within a transaction (i.e. plain JDBC code working with the same DataSource)"; "To be able to register a DataSource's Connection for plain JDBC code, this instance needs to be aware of the DataSource (`setDataSource(DataSource)`)"; "This transaction manager will autodetect the DataSource used as the connection factory of the EntityManagerFactory, so you usually don't need to explicitly specify the 'dataSource' property"; and "Note that this requires a vendor-specific `JpaDialect` to be configured" +- https://docs.spring.io/spring-framework/docs/current/javadoc-api/org/springframework/jdbc/datasource/DataSourceUtils.html — `applyTimeout` "Apply the specified timeout - overridden by the current transaction timeout, if any - to the given JDBC Statement object"; `applyTransactionTimeout` "Apply the current transaction timeout, **if any**, to the given JDBC Statement object" — the "if any" is the silent branch +- https://github.com/spring-projects/spring-framework/blob/v6.2.0/spring-orm/src/main/java/org/springframework/orm/jpa/JpaTransactionManager.java — `doBegin` sets `conHolder.setTimeoutInSeconds(timeoutToUse)` only inside `if (getDataSource() != null)` and only when `getJpaDialect().getJdbcConnection(em, …)` returned non-null, otherwise logging "Not exposing JPA transaction … because JpaDialect … does not support JDBC Connection retrieval". `DefaultJpaDialect.getJdbcConnection` returns `null`. Read at tag v6.2.0 +- https://github.com/spring-projects/spring-framework/blob/v6.2.0/spring-jdbc/src/main/java/org/springframework/jdbc/support/SQLStateSQLExceptionTranslator.java — `DATA_ACCESS_RESOURCE_FAILURE_CODES` is `Set.of("08", "53", "54", "57", "58")` and class `57` returns `new DataAccessResourceFailureException(...)` with no timeout special-case; the only `QueryTimeoutException` route is `ex.getClass().getName().contains("Timeout")` (commented "For MySQL"). Verified 2026-08-10 across tags: `"57014".equals(sqlState)` is absent in v5.3.31, v6.0.0, v6.2.0, v6.2.1, v6.2.3, v6.2.5 and v6.2.8, and present in v7.0.0 and `main` (as `indicatesQueryTimeout`, documented "with SQL state 57014 as a specific indication") +- `SQLExceptionSubclassTranslator` (same tag) maps `ex instanceof SQLTimeoutException` to `QueryTimeoutException` and constructs `setFallbackTranslator(new SQLStateSQLExceptionTranslator())`; `JdbcAccessor` documents it as the default "as of 6.0" unless a user-provided `sql-error-codes.xml` is on the classpath. PgJDBC's `PSQLException extends SQLException` and its `PSQLState.QUERY_CANCELED` is `"57014"`, so the subclass branch does not match and the state fallback decides — read from https://github.com/pgjdbc/pgjdbc/blob/master/pgjdbc/src/main/java/org/postgresql/util/PSQLException.java and `PSQLState.java` +- https://github.com/hibernate/hibernate-orm/blob/main/hibernate-core/src/main/java/org/hibernate/dialect/PostgreSQLDialect.java — SQLState `"57014"` maps to `org.hibernate.QueryTimeoutException`, which `HibernateJpaDialect` converts to `org.springframework.dao.QueryTimeoutException` (https://github.com/spring-projects/spring-framework/blob/v6.2.0/spring-orm/src/main/java/org/springframework/orm/jpa/vendor/HibernateJpaDialect.java) — the JPA half of the split in step 4 +- https://mybatis.org/mybatis-3/configuration.html — `defaultStatementTimeout` "Sets the number of seconds the driver will wait for a response from the database"; valid values "Any positive integer"; default "Not Set (null)" — so an unset value means MyBatis applies no bound of its own +- https://github.com/mybatis/mybatis-3/blob/273ec6508c348513456cb2cc48f943b28196a56d/src/main/java/org/apache/ibatis/executor/statement/BaseStatementHandler.java — `setStatementTimeout` resolves `mappedStatement.getTimeout()` first, falls back to `configuration.getDefaultStatementTimeout()`, calls `stmt.setQueryTimeout(queryTimeout)` when either is non-null, then calls `StatementUtil.applyTransactionTimeout(stmt, queryTimeout, transactionTimeout)`, which — when a transaction timeout exists — lowers the statement to it whenever `queryTimeout` is `null`, is `0` (JDBC's "no limit" sentinel), or is larger than the transaction's remaining time. The bound therefore exists with no `ConnectionHolder` involved — the opposite of the `DataSourceUtils.applyTimeout` path above +- https://github.com/postgres/postgres/blob/083ac033419f690758508e08c1736089384bbee8/src/backend/tcop/postgres.c — `ProcessInterrupts` emits `"canceling statement due to lock timeout"`, `"canceling statement due to statement timeout"`, `"canceling autovacuum task"` and, as the fall-through, `"canceling statement due to user request"`; all but the lock-timeout case use `ERRCODE_QUERY_CANCELED` (57014), so the message string is what distinguishes a server-side statement timeout from an external cancel +- Field measurement 2026-08-11 (same production service, p6spy JDBC timing): the MyBatis path cancelled at 120,010 ms against a configured `defaultStatementTimeout` of 120 s, and a repo-wide grep found no other layer configured at 120 s or 10 s — the duration-to-configured-value match is what attributed each cancellation to its layer +- Field measurement 2026-08-10 (production endpoint, p6spy JDBC timing, `@Transactional(readOnly = true, timeout = 10)` declared, no server-side `statement_timeout`): the Hibernate path was cancelled at 10,012 ms and surfaced as HTTP 400; two raw `JdbcTemplate` calls on the same endpoint ran 151,558 ms and 163,489 ms and surfaced as HTTP 500. Over 30 days, 11 recorded errors matched the per-path status split with no exceptions diff --git a/wiki/backend/python/index.md b/wiki/backend/python/index.md index c0473f2..762a1f5 100644 --- a/wiki/backend/python/index.md +++ b/wiki/backend/python/index.md @@ -24,9 +24,16 @@ Match your situation to a "load when" line; load only matching pages. |------|-----------| | [app-servers-and-workers](serving/app-servers-and-workers.md) | Deploying a Python web app behind gunicorn/uvicorn; choosing worker count, worker class (sync/gthread/ASGI), or worker timeout; requests queue or time out while CPU sits idle; workers killed mid-request; worker memory growth and preload/max_requests recycling decisions | +## packaging + +| Page | Load when | +|------|-----------| +| [data-files-and-install-paths](packaging/data-files-and-install-paths.md) | A Python package reads non-code files (grammars, templates, KBs) located via `__file__`-relative paths; an installed console script cannot find a data file that exists in the repo; deciding how data files get into the wheel and how code resolves them after `pip install .` | + ## language | Page | Load when | |------|-----------| | [mutable-state-traps](language/mutable-state-traps.md) | State persists or leaks across calls/requests in a long-lived Python process — one user's data appears for another, values "remembered" between calls; loop-built callbacks all use the last value; reviewing function signatures (mutable defaults), class bodies (class attributes), or module-level objects for hidden sharing; choosing contextvars vs thread-locals for request context | | [bytecode-cache-staleness](language/bytecode-cache-staleness.md) | A script or harness rewrites `.py` files and re-runs them in a loop (mutation testing, edit/test/revert, codegen check, bisect) and the result stops tracking what is on disk — a revert that `git diff` reports clean still fails, or an injected change has no effect; choosing between clearing `__pycache__`, refreshing mtime, and hash-based `.pyc` (PEP 552); designing byte-length-preserving mutations | +| [default-encoding-in-text-io](language/default-encoding-in-text-io.md) | Python opens a text file without `encoding=` (`open`, `Path.read_text`, `subprocess` text mode) and you are adding the argument or writing the regression test that keeps it there; a file-writing bug reproduces on Windows, a `LANG=C` container, or a cp949/cp932 desktop but not on your machine; choosing a test discriminator that does not depend on the runner's locale | diff --git a/wiki/backend/python/language/bytecode-cache-staleness.md b/wiki/backend/python/language/bytecode-cache-staleness.md index c911edd..ccdff9c 100644 --- a/wiki/backend/python/language/bytecode-cache-staleness.md +++ b/wiki/backend/python/language/bytecode-cache-staleness.md @@ -8,8 +8,10 @@ sources: - https://docs.python.org/3/reference/import.html - https://peps.python.org/pep-0552/ - https://docs.python.org/3/library/py_compile.html -last_verified: 2026-08-04 -related: [testing-quality-harness-reverse-controls, testing-quality-tests-that-cannot-fail, backend-python-language-mutable-state-traps] + - https://docs.python.org/3/library/shutil.html + - https://docs.python.org/3/using/cmdline.html +last_verified: 2026-08-11 +related: [backend-python-language-default-encoding-in-text-io, testing-quality-harness-reverse-controls, testing-quality-tests-that-cannot-fail, backend-python-language-mutable-state-traps] --- # Edited Python Source the Interpreter Keeps Ignoring @@ -55,7 +57,9 @@ no effect at all. | The mutation is the thing that vanished (injected change has no effect) and the revert looks fine | Same mechanism, opposite direction: the cache predates both writes. Clear `__pycache__` and re-inject, then confirm the mutation *does* change behavior before scoring it as "caught" ([testing-quality-harness-reverse-controls]) | | The harness reports every mutant caught | Verify one mutation reaches the interpreter by hand first — a stale cache that pins the *original* bytecode makes every mutant look survived, and one that pins a *mutant* makes every later case look caught | | Writes are driven by a tool that preserves mtime (`rsync -t`, archive extraction, `git checkout` of an unchanged blob, `touch -t` in a script) | The second-granularity race becomes a certainty rather than a race; use hash-based `.pyc` or clear the cache unconditionally | -| The tree is read-only or `PYTHONDONTWRITEBYTECODE` is set | No `.pyc` is written, so this failure cannot occur — and the harness pays a recompile per run | +| The harness backs the file up and restores it with `shutil.copy2` | `copy2` "also attempts to preserve file metadata" via `copystat`, so the restore stamps the *original* mtime back — the same certainty as the row above, reached through the idiomatic backup/restore call. Apply that row's remedy (hash-based `.pyc`, or clear `__pycache__` unconditionally), and switch the restore to `shutil.copyfile` ("no metadata") or `shutil.copy` ("the file's creation and modification times, is not preserved") so the write at least stops re-stamping the old timestamp. A bare `os.utime(path, None)` after the restore is not sufficient on its own: it sets mtime to *now*, which collides again whenever the cached compile happened in the same second | +| The tree is read-only, or the harness runs under `-B` / `PYTHONDONTWRITEBYTECODE` | Both settings govern writing only — Python "won't try to write `.pyc` files on the import of source modules" — so a `.pyc` left on disk by an earlier run is still validated and reused, and this failure still occurs. Purge `__pycache__` once before the run and keep `-B` set for the rest of it: with nothing cached and nothing written, every iteration recompiles from source | +| The harness re-imports the mutated file through a fresh `importlib.util.spec_from_file_location` + `module_from_spec` + `exec_module` each iteration | A fresh spec and module object bypass `sys.modules`, not the on-disk cache: the source loader still validates `__pycache__` and reuses it. Purge the cache (or compile with hash-based invalidation) between iterations, or read the file yourself and `exec(compile(text, path, "exec"), ns)`, which consults no cache | | The stale module was already imported in a long-lived process | Clearing `__pycache__` does not help; the module object is in `sys.modules` and only a fresh process (or an explicit reload) picks the change up | | An installed package ships `.pyc` files without sources | The unchecked-hash variant is assumed valid whenever it exists; edits to a co-located source are never consulted | @@ -73,4 +77,8 @@ no effect at all. - https://docs.python.org/3/reference/import.html — "By default, Python does this by storing the source's last-modified timestamp and size in the cache file when writing it"; "At runtime, the import system then validates the cache file by checking the stored metadata in the cache file against the source's metadata"; hash-based `.pyc` files store "a hash of the source file's contents rather than its metadata", in checked and unchecked variants, overridable with `--check-hash-based-pycs` - https://peps.python.org/pep-0552/ — hash-based `.pyc` invalidation, added in Python 3.7, as the deterministic alternative to timestamp+size - https://docs.python.org/3/library/py_compile.html — `PycInvalidationMode` selects timestamp, checked-hash, or unchecked-hash invalidation when compiling +- https://docs.python.org/3/using/cmdline.html — `-B`: "If given, Python won't try to write `.pyc` files on the import of source modules"; `PYTHONDONTWRITEBYTECODE` "is equivalent to specifying the `-B` option" — both describe writing only, so neither stops an existing `.pyc` from being read +- https://docs.python.org/3/library/shutil.html — `copyfile` copies "the contents (no metadata)"; `copy` copies data and permission mode and "Other metadata, like the file's creation and modification times, is not preserved"; `copy2` is "Identical to `copy()` except that `copy2()` also attempts to preserve file metadata" and "uses `copystat()` to copy the file metadata" — so `copy2` is the mtime-restoring member of the family +- Field reproduction 2026-08-11 (batch mutation harness, one process, byte-length-preserving mutation of a numeric literal, restore via `shutil.copy2`): three consecutive mutants scored GREEN in the batch and the third scored RED when run alone; printing the mutated constant from a fresh subprocess showed all three runs loading the *first* mutant's value. Purging `__pycache__` and calling `os.utime(path, None)` between iterations flipped the third to RED while a no-op control mutation stayed GREEN +- Field reproduction 2026-08-11 (Python 3.14.6, macOS, fresh-spec import path): `mod.py` holding `VERSION = "3.1.0"` was pinned to a fixed mtime and loaded via `spec_from_file_location` + `module_from_spec` + `exec_module`, which wrote `__pycache__/mod.cpython-314.pyc`. Rewriting the file to the same-length `"3.1.1"` and re-pinning the same mtime, the identical loader printed `3.1.0`; re-running under `python3 -B` with that `.pyc` still present also printed `3.1.0`. A fresh spec and `-B` each leave stale bytecode in play - Field reproduction 2026-08-04 (Python 3.14.6, macOS): with `mod.py` pinned to a fixed mtime via `touch -t` and every revision exactly 18 bytes, compiling `VERSION = "3.1.1"` and then reverting the file to `VERSION = "3.1.0"` left `import mod` reporting `3.1.1` — reverted source, mutant bytecode. Deleting `__pycache__` returned `3.1.0`; a bare `touch mod.py` (mtime bumped, cache left in place) also returned `3.1.0`. The `.pyc` header decoded to `flags=0` (timestamp invalidation) with the source's exact mtime and `size=18` diff --git a/wiki/backend/python/language/default-encoding-in-text-io.md b/wiki/backend/python/language/default-encoding-in-text-io.md new file mode 100644 index 0000000..e3b6917 --- /dev/null +++ b/wiki/backend/python/language/default-encoding-in-text-io.md @@ -0,0 +1,107 @@ +--- +id: backend-python-language-default-encoding-in-text-io +domain: backend +category: language +applies_to: [python] +confidence: verified +sources: + - https://peps.python.org/pep-0597/ + - https://peps.python.org/pep-0686/ + - https://docs.python.org/3/library/functions.html +last_verified: 2026-08-07 +related: + [ + backend-python-language-bytecode-cache-staleness, + platforms-environment-timezone-and-locale, + testing-quality-tests-that-cannot-fail, + testing-quality-minimum-case-set, + ] +--- + +# Text I/O Whose Encoding Comes from the Machine's Locale + +## When this applies + +Python code opens a text file without `encoding=` — `open(p)`, `open(p, "w")`, +`Path.read_text()`, `csv`/`json` wrappers built on them — and you are adding the +argument, or writing the regression test that keeps it there. Also when a +file-writing bug reproduces on one machine (Windows, a `LANG=C` container, a +cp949/cp932 desktop) and not on yours. + +Timezone and locale as hidden inputs across dates and text → +[platforms-environment-timezone-and-locale]. + +## Do this + +1. **Pass `encoding=` at every text-mode call site.** The default is the + machine's: "The default encoding is platform dependent (whatever + `locale.getencoding()` returns)". Choose the value from what the file is: + +| The file is | Pass | +| --------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- | +| A format with a defined encoding (JSON, TOML, YAML, Markdown, source) | `encoding="utf-8"` | +| Written and read only by this program | `encoding="utf-8"` | +| Produced by a tool bound to the OS console encoding, deliberately | `encoding=locale.getencoding()`, stated explicitly so the dependency is visible | +| Raw bytes | Binary mode with no `encoding` — "For reading and writing raw bytes use binary mode and leave _encoding_ unspecified" | + +2. **Make the regression test run the real entry point under the interpreter's + own diagnostic, and assert zero warnings naming the file you fixed:** + + ```sh + python3 -X warn_default_encoding -W always::EncodingWarning + ``` + + `EncodingWarning` "is emitted when the `encoding` argument to `open()` is + omitted and the default locale-specific encoding is used", and the flag (or + `PYTHONWARNDEFAULTENCODING`) is what enables it. Filter the captured stderr + to the file under test by name, so unfixed call sites elsewhere in the + codebase do not redden this test. + +3. **Assert on the warning lines, not on the produced bytes.** The warning is + emitted at the call site regardless of what the locale happens to be, so it + is the same verdict on your laptop and in CI. + +4. **Prove the check reddens before trusting it.** Without `-X + warn_default_encoding` the warning is silent, so a runner that drops the flag + reports green on the reintroduced defect and looks identical to a pass. Seed a + deliberately unencoded `open()` in the file under test, require red, then + restore ([testing-quality-tests-that-cannot-fail]). + +5. **Widen the flag from the one test invocation to the whole CI run once every + call site is clean**, so a new omission is caught where it is written rather + than at the next locale change. Until then the filter in step 2 is what keeps + the unfixed sites from reddening this test. + +6. **Keep a value assertion for the encodings you set explicitly.** + `EncodingWarning` fires only on an *omitted* argument, so it says nothing about + `encoding="latin-1"` or a deliberate `encoding=locale.getencoding()`. For those + call sites, assert the bytes the file should contain, and run that assertion + under a non-UTF-8 locale (`LANG=C`, or a cp949/cp932 job) where a wrong value + changes the output. + +## Edge cases + +| Case | Then | +| ---------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| The entry point is a library function, not a script | Run it through a one-line driver under the same flags; the warning is attributed to the frame that called `open()`, so the driver's own lines do not mask it | +| A dependency emits `EncodingWarning` from its own files | Filter by filename as in step 2 and record the dependency in the test's name, so the filter states what it is excluding | +| The code targets Python 3.15 or later, where UTF-8 mode is on by default (PEP 686) | Keep the explicit `encoding=`: the argument states the file's contract and is what makes the call correct under an inherited `PYTHONUTF8=0` or an older runtime | +| Running under `PYTHONUTF8=1` / UTF-8 mode already | The warning still fires on the omitted argument, so the test keeps working; the mode changes the value used, not whether the argument was passed | +| `subprocess` output is being decoded | The same default applies to its text mode — pass `encoding="utf-8"` there, and include it in the call-site sweep | +| The harness rewrites the file between runs to seed the missing-`encoding` mutation | Clear the bytecode cache between iterations ([backend-python-language-bytecode-cache-staleness]) | + +## Instead of + +| If you are about to | Do this instead | Why | +| --------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| Prove an omitted-`encoding` fix with a round-trip assertion alone (write non-ASCII text, read it back, compare) | Assert zero `EncodingWarning` lines naming the file, under `-X warn_default_encoding`, and keep the round-trip for the explicitly-set encodings (step 6) | On a UTF-8 locale the encoded bytes are identical with and without the argument, so the round-trip is green on the defect — it discriminates only on a runner whose locale encoding is not UTF-8, which is not the default on macOS or on most Linux CI images | +| Assert the output file's declared charset (``, an XML declaration) | Assert the warning count | A declaration is a literal in the template — it is written correctly by code that encoded the body wrongly | +| Set `LANG`/`PYTHONUTF8` in the test environment to make the behavior deterministic | Fix the call sites and assert the warning | Pinning the environment makes the test pass by removing the input the defect depends on, so the defect ships and fails on the machines that do not inherit that environment | +| Read "it works on macOS and Linux" as evidence the encoding is right | Run the warning check | PEP 686: "many Python developers using Unix forget that the default encoding is platform dependent … Inconsistent default encoding causes many bugs"; "this change mostly affects Windows users" | + +## Sources + +- https://peps.python.org/pep-0597/ — `EncodingWarning` "is emitted when the `encoding` argument to `open()` is omitted and the default locale-specific encoding is used"; "The `-X warn_default_encoding` option and the `PYTHONWARNDEFAULTENCODING` environment variable are added. They are used to enable `EncodingWarning`"; "When the flag is set, `io.TextIOWrapper()`, `open()` and other modules using them will emit `EncodingWarning` when the `encoding` argument is omitted"; "Developers using macOS or Linux may forget that the default encoding is not always UTF-8" +- https://peps.python.org/pep-0686/ — enabling UTF-8 mode by default targets Python 3.15; "many Python developers using Unix forget that the default encoding is platform dependent. They omit to specify `encoding="utf-8"` … Inconsistent default encoding causes many bugs"; "Most Unix systems use UTF-8 locale … So this change mostly affects Windows users" +- https://docs.python.org/3/library/functions.html — `open()`: "The default encoding is platform dependent (whatever `locale.getencoding()` returns)"; "In text mode, if _encoding_ is not specified the encoding used is platform-dependent"; "For reading and writing raw bytes use binary mode and leave _encoding_ unspecified" +- Reproduction 2026-08-07 (CPython 3.14.6, macOS, `locale.getpreferredencoding(False) == 'UTF-8'`): a script with one `open(p, "w")` and one `open(p, "w", encoding="utf-8")` produced byte-identical output — a round-trip assertion cannot distinguish them. `python3 -X warn_default_encoding -W always::EncodingWarning script.py out.txt` emitted exactly one line, naming the unencoded call by file and line number; the same run without the flag emitted nothing diff --git a/wiki/backend/python/packaging/data-files-and-install-paths.md b/wiki/backend/python/packaging/data-files-and-install-paths.md new file mode 100644 index 0000000..13ddfa7 --- /dev/null +++ b/wiki/backend/python/packaging/data-files-and-install-paths.md @@ -0,0 +1,65 @@ +--- +id: backend-python-packaging-data-files-and-install-paths +domain: backend +category: packaging +applies_to: [python] +confidence: verified +sources: + - https://setuptools.pypa.io/en/latest/userguide/datafiles.html + - https://docs.python.org/3/library/importlib.resources.html +last_verified: 2026-08-14 +related: [testing-data-test-data-and-isolation] +--- + +# A Python Package That Reads Data Files Shipped Next to Its Code + +## When this applies + +A Python package reads non-code files at runtime — grammar definitions, dialect +files, templates, knowledge bases — and locates them with `__file__`-relative +paths (`Path(__file__).parent / ...` or `.parents[N]`); or an installed console +script cannot find a data file that exists in the repository. + +## Do this + +1. **Declare the data files as package data** so the build backend puts them in + the wheel: `include_package_data = True` (files matched by `MANIFEST.in` or + tracked by VCS) or explicit `package_data` / `tool.setuptools.package-data` + globs. A file the wheel does not contain cannot be found by any path logic + after install. +2. **Resolve them by package name, not filesystem position**: + `importlib.resources.files("mypkg.data").joinpath("grammar.mlir").read_text()`. + When a real on-disk path is required (a subprocess takes a filename), wrap it + in `as_file()` — the context manager extracts from a zip when necessary and + removes the extraction on exit. +3. **Keep every path anchor inside the package.** A `.parents[N]` walk that + escapes the package directory resolves into the venv's internals after a + non-editable install — the repo layout above the package does not exist in + `site-packages`. +4. **Verify against a real non-editable install, not the checkout**: build and + install into a scratch venv with `pip install .` (no `-e`), then run the + entry point from a directory outside the repo. Editable installs and + `PYTHONPATH` runs anchor `__file__` in the repo, so they pass even when the + wheel ships zero data files. + +## Edge cases + +| Case | Then | +|------|------| +| The data file must be handed to an external tool as a path | `with as_file(files("pkg.data") / name) as p:` and consume `p` inside the block — a zip extraction is deleted when the block exits | +| Data lives outside any package directory (repo-root `data/`) | Move it under an importable package — setuptools includes data files per package, and `files()` resolves by importable name | +| Python < 3.9 must run the code | Use the `importlib-resources` backport, which provides the same `files()` API | +| Every existing test runs from the repo checkout | Add one job that installs the built wheel into a clean venv and smoke-tests the CLI from outside the repo — repo-anchored resolution failures are invisible to in-repo tests | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Compute a data path as `Path(__file__).resolve().parents[N] / "data"` | Resolve via `importlib.resources.files()` on the owning package | `__file__` anchors wherever the module was imported from; after `pip install .` that is `site-packages`, and the `parents[N]` walk lands inside the venv where the repo's directories do not exist | +| Prove packaging works by running the suite in the checkout | Install the wheel into a scratch venv and run the entry point from outside the repo | Editable and `PYTHONPATH` runs resolve `__file__` to the repo, masking a wheel that ships no data files | + +## Sources + +- https://setuptools.pypa.io/en/latest/userguide/datafiles.html — "It is strongly recommended that, if you are using data files, you should use `importlib.resources` to access them"; `include_package_data` / `package_data` control what the wheel contains; `__file__` manipulation "isn't compatible with PEP 302-based import hooks, including importing from zip files" +- https://docs.python.org/3/library/importlib.resources.html — packages and resources "do not have to exist as physical files and directories on the file system"; `as_file()` yields a `pathlib.Path` and cleans up any temporary extraction on exit +- Field reproduction 2026-08-14 (linkly v0.4.0): installed `lnpl build` failed rc=4 — the `__file__`-anchored grammar path resolved to `.venv/lib/python3.13/mlir/lnpl.irdl.mlir`, which the wheel never shipped (`impl/lnpl/backend.py:63-64`), while `PYTHONPATH=impl` runs of the same source built fine diff --git a/wiki/databases/data-survey/audit-columns-as-update-evidence.md b/wiki/databases/data-survey/audit-columns-as-update-evidence.md new file mode 100644 index 0000000..6687da3 --- /dev/null +++ b/wiki/databases/data-survey/audit-columns-as-update-evidence.md @@ -0,0 +1,87 @@ +--- +id: databases-data-survey-audit-columns-as-update-evidence +domain: databases +category: data-survey +applies_to: [postgresql, mysql, general] +confidence: verified +sources: + - https://github.com/spring-projects/spring-data-jpa/blob/main/spring-data-jpa/src/main/java/org/springframework/data/jpa/domain/support/AuditingEntityListener.java + - https://docs.hibernate.org/orm/6.6/querylanguage/html_single/Hibernate_Query_Language.html + - https://github.com/hibernate/hibernate-orm/blob/6.6/hibernate-core/src/main/java/org/hibernate/event/internal/DefaultFlushEntityEventListener.java + - https://www.postgresql.org/docs/current/sql-createtrigger.html +last_verified: 2026-08-13 +related: [databases-data-survey-surveying-live-data-for-a-rule, databases-schema-design-nullability-and-defaults, backend-java-jpa-not-null-check-and-lifecycle-callbacks] +--- + +# Audit Columns as Evidence About a Row's Update History + +## When this applies + +You are surveying live rows and about to turn an audit column into a behavioural +claim: `update_dt`/`updated_at`/`modified_by` is NULL (or unchanged) on every row of +interest and you read that as "these rows were never modified", "this feature is +unused", or "the bad values came in at insert and nothing touched them since". Also +when an incident investigation needs to know whether writes to those rows were +*attempted*. + +Deriving a mapping or enum rule from a survey → [databases-data-survey-surveying-live-data-for-a-rule]. + +## Do this + +1. **Identify what writes the column before reading it as history**, and bound the + claim to that writer: + +| Written by | Records | A NULL therefore means | +|------------|---------|------------------------| +| ORM lifecycle callback (`@PreUpdate`; Spring Data's `AuditingEntityListener.touchForUpdate` is `@PreUpdate`) | Only updates that reached the entity's flush action | No *successful, entity-level* update ran | +| Application code assigning the field | Only the code paths that assign it | No update through those paths | +| DB trigger (`BEFORE UPDATE … FOR EACH ROW`) or a generated/`ON UPDATE` column | Every statement the database executed on the row, whatever the client | The row's columns were not updated | + +2. **State the bounded claim in the deliverable**: "no update ran through the path that + stamps this column, as of " — not "never updated". Name the writer you found + in directive 1 alongside the count. +3. **Enumerate the write paths that leave the column untouched** before concluding + anything, and check each against the code: + +| Path | Why the column stays as it was | +|------|-------------------------------| +| The update failed pre-flush | Validation that runs before the update action is scheduled — e.g. Hibernate's not-null check → [backend-java-jpa-not-null-check-and-lifecycle-callbacks] — throws before any `@PreUpdate` listener runs | +| Bulk JPQL/HQL `update`, or native SQL | "The effect of an `update` or `delete` statement is not reflected in the persistence context": no entity flush, so no callback and no `@Version` bump unless the statement is `versioned` | +| Another service, migration, or manual SQL writes the table | It never loaded the entity, so the ORM's auditing was never in the path | +| The transaction rolled back after the callback set the field | The in-memory stamp is discarded with the transaction | + +4. **Judge update history on an independent axis** — a history/audit table, the + application logs for the writing endpoint, CDC/WAL, or a DB-side trigger installed + going forward. Record which axis the conclusion rests on. +5. **Require one positive control before acting on the absence.** Find at least one row + whose audit column *is* set by the same writer (or a test that exercises it). Absent + that, the NULLs are equally explained by "the writer never worked here", and any + decision built on them (backfill, drop the column, close the ticket as "unused") is + resting on an unmeasured mechanism. + +## Edge cases + +| Case | Then | +|------|------| +| Only old rows are NULL | The auditing was added later — find the migration or the earliest non-null value and read NULL as "before instrumentation", not as behaviour | +| `updated_at` equals `created_at` on every row | The writer stamps both at insert; equality is evidence of no update only if you have confirmed the update path stamps it (directive 5) | +| Every non-null value shares one timestamp | A backfill or bulk migration wrote them; those rows carry no per-row update history | +| The column is `NOT NULL DEFAULT now()` | It cannot distinguish "inserted" from "updated" at all — pair it with a separate insert timestamp or a history table | +| The claim needed is "did anyone *read*/attempt this" | Audit columns cannot answer it in any configuration; go to application logs or the DB's statement/audit logging | +| Rows are soft-deleted | The delete may run as an update through a different path than the business update → [databases-schema-design-soft-delete] | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Conclude "never modified" from all-NULL audit columns | State the bounded claim (directive 2), then confirm on an independent axis (directive 4) | The column records successes of one path; failed, bulk, and out-of-band writes leave it untouched | +| Treat the audit column as the failure timeline in an incident | Use application logs or a history table for attempts, and keep the audit column for confirmed successes | The investigation's subject is the failing write, which is exactly the event the column cannot record | +| Add `@PreUpdate` auditing to close the gap you just found | Add a DB trigger (or generated column) when the requirement is "every statement, whatever the client" | Callback auditing is bypassed by bulk DML, native SQL, and other services by construction | + +## Sources + +- https://github.com/spring-projects/spring-data-jpa/blob/main/spring-data-jpa/src/main/java/org/springframework/data/jpa/domain/support/AuditingEntityListener.java — `AuditingEntityListener.touchForCreate` is annotated `@PrePersist` and `touchForUpdate` `@PreUpdate`, so `@LastModifiedDate` is written by a JPA lifecycle callback +- https://docs.hibernate.org/orm/6.6/querylanguage/html_single/Hibernate_Query_Language.html — mutation statements: "The effect of an `update` or `delete` statement is not reflected in the persistence context, nor in the state of entity objects held in memory at the time the statement is executed"; "It's the responsibility of the client program to maintain synchronization of state held in memory with the database"; `update` leaves `@Version` attributes alone unless the `versioned` keyword is used +- https://github.com/hibernate/hibernate-orm/blob/6.6/hibernate-core/src/main/java/org/hibernate/event/internal/DefaultFlushEntityEventListener.java — `scheduleUpdate` runs the nullability check before `EntityUpdateAction` is queued, so a pre-flush validation failure precedes every `@PreUpdate` listener +- https://www.postgresql.org/docs/current/sql-createtrigger.html — "A trigger that is marked `FOR EACH ROW` is called once for every row that the operation modifies"; the trigger is defined on the relation, not in a client, which is what makes it the axis independent of the application's write path +- Field observation 2026-08-13 (PostgreSQL, `manage.building_tenant_floor_info`): all 574 rows with a NULL business code also had `update_dt` NULL, which read as "never edited"; those rows' UPDATEs were in fact failing in Hibernate's `Nullability` check before `EntityUpdateAction` was created, so the `@PreUpdate` stamp never ran. The same repository also updates a table by bulk JPQL without setting `updateDt`, a second path with the same signature diff --git a/wiki/databases/data-survey/catalog-statistics-as-current-state.md b/wiki/databases/data-survey/catalog-statistics-as-current-state.md new file mode 100644 index 0000000..5d2a3fc --- /dev/null +++ b/wiki/databases/data-survey/catalog-statistics-as-current-state.md @@ -0,0 +1,111 @@ +--- +id: databases-data-survey-catalog-statistics-as-current-state +domain: databases +category: data-survey +applies_to: [postgresql] +confidence: verified +sources: + - https://www.postgresql.org/docs/current/catalog-pg-class.html + - https://www.postgresql.org/docs/current/functions-admin.html + - https://www.postgresql.org/docs/current/runtime-config-preset.html + - https://www.postgresql.org/docs/current/sql-analyze.html + - https://www.postgresql.org/docs/current/monitoring-stats.html + - https://www.postgresql.org/docs/current/ddl-system-columns.html + - https://www.postgresql.org/docs/current/datatype-oid.html + - https://www.postgresql.org/docs/release/14.0/ +last_verified: 2026-08-12 +related: [databases-data-survey-surveying-live-data-for-a-rule, databases-operations-autovacuum-and-wraparound, databases-query-optimization-existence-and-count-checks, databases-query-optimization-keyset-pagination] +--- + +# Catalog Statistics as Evidence About a Table's Current Contents + +## When this applies + +You need to know what the newest rows of a large PostgreSQL table are — the latest +batch date, whether an ingest still runs, which values arrived this month — and the +column carrying that answer has no index, so a full scan is too expensive. You reach +for the catalog instead: `pg_class.relpages`/`reltuples`, `pg_stats` most-common +values, or a `ctid` range scan over the last blocks. Also when a cheap probe came +back with only old values and you are about to record "no recent data". + +Deriving a mapping or enum rule from a survey → [databases-data-survey-surveying-live-data-for-a-rule]. + +## Do this + +1. **Take the table's physical end from `pg_relation_size`, not from `relpages`.** + `relpages` is "only an estimate used by the planner" that is "updated by `VACUUM`, + `ANALYZE`, and a few DDL commands" — it is a snapshot of the last maintenance run, + so every block appended since is invisible in it. `pg_relation_size(rel)` reads the + main fork's actual byte size: + + ```sql + SELECT pg_relation_size('external_data.t') + / current_setting('block_size')::int AS last_block; + ``` + +2. **Divide by `current_setting('block_size')`, not by a literal `8192`.** The + preset reports "the size of a disk block ... determined by the value of `BLCKSZ` + when building the server"; 8192 is the default, not a guarantee. + +3. **Choose the probe by which question you are answering:** + +| Question | Probe | +|----------|-------| +| What values exist near the physical end? | `SELECT DISTINCT col FROM t WHERE ctid > '(,0)'::tid` — aggregate over the range, so every block in it contributes | +| Does any row newer than X exist? | `SELECT max(col) FROM t WHERE ctid > '(,0)'::tid`, widening N until the answer stops moving | +| How many rows sit in this range? | `count(*)` over the same predicate — a zero here is unambiguous, an empty `DISTINCT` is not | + +4. **When you bound a range probe with `LIMIT n`, read the result as "the first n + rows at the start of the range", not as "the range's contents".** The scan walks + the range in ascending block order, so `LIMIT` truncates at its oldest end — the + newest blocks are exactly what it drops. + +5. **Before citing `pg_stats` (most-common values, histogram) as the value set, read + `last_analyze`, `last_autoanalyze`, and `n_mod_since_analyze` from + `pg_stat_all_tables`.** ANALYZE "takes a random sample of the table contents, + rather than examining every row" and its output is "only approximate"; when the + last analyze is old or `NULL`, values that arrived since are absent from the MCV + list by construction, not by absence in the table. + +6. **Cross-check `reltuples` against `n_live_tup` and treat a gap as unmeasured + inflow.** `reltuples` moves only at VACUUM/ANALYZE/DDL, while `n_live_tup` is the + cumulative-statistics estimate; the difference is the signal that the catalog view + of the table is behind the table. + +7. **Report which probe produced the answer, with the block range and the analyze + timestamps you read.** A tail-scan answer is a statement about the blocks scanned; + without the range it reads as a statement about the table. + +## Edge cases + +| Case | Then | +|------|------| +| The table takes UPDATEs or DELETEs, not append-only inserts | Physical order stops tracking insert order — a row's `ctid` "will change if it is updated or moved by `VACUUM FULL`", and freed space is refilled by later writes. Confirm the append-mostly assumption (`n_tup_upd`/`n_tup_del` near zero) before reading the tail as "newest" | +| The server is PostgreSQL 13 or older | Range predicates on `ctid` fall back to a sequential scan; efficient TID range scanning arrived in 14 ("Previously a sequential scan was required for non-equality `TID` specifications"). Bound the cost another way, or accept the scan | +| `reltuples` is `-1` | The table "has never yet been vacuumed or analyzed" and the row count is unknown — no catalog-derived count exists to compare against | +| The last blocks are empty of visible rows | The tail can hold dead tuples or free space; widen the range downward rather than concluding the table stopped receiving rows | +| The size you need includes TOASTed values | `pg_relation_size` with one argument returns the main fork only; `pg_table_size` adds TOAST, FSM, and visibility map. Block arithmetic for `ctid` uses the main fork | +| You may run maintenance on the table | `ANALYZE ` refreshes the statistics and makes `pg_stats`/`relpages` answer the question directly — take this path when the table is not so large that the analyze cost matters | +| The probe runs against production | Keep every range bounded by `ctid` and read `EXPLAIN` before executing; an unbounded probe on the same column is the full scan you were avoiding ([databases-query-optimization-reading-execution-plans]) | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Start a tail scan at `relpages` | Start at `pg_relation_size(rel)/current_setting('block_size')::int` | `relpages` is the planner's estimate from the last VACUUM/ANALYZE; blocks appended since sit above it and are never scanned | +| Read "the MCV list's newest value is April" as "nothing arrived after April" | Read `last_analyze`/`last_autoanalyze` first, and re-probe the heap when they are old or `NULL` | The MCV list is a sample from the moment of the last analyze, so recency of data and recency of statistics are different facts | +| Bound a `ctid` range probe with `LIMIT 20` and read the values as the range's | Aggregate the range (`DISTINCT`, `max`, `count`) or narrow it to the last blocks | The scan returns the range's oldest rows first, so `LIMIT` yields the end you were not asking about | +| Hard-code `/ 8192` in the block arithmetic | Read `current_setting('block_size')` | `BLCKSZ` is a build-time choice; the arithmetic silently points at the wrong block on a non-default build | +| Conclude "the ingest stopped" from one cheap catalog probe | Name the probe and its range in the finding, then confirm with a second probe of a different kind | Catalog estimates and heap contents are different sources; agreement between two of them is what makes the conclusion an observation | + +## Sources + +- https://www.postgresql.org/docs/current/catalog-pg-class.html — `relpages`: "Size of the on-disk representation of this table in pages (of size `BLCKSZ`). This is only an estimate used by the planner. It is updated by `VACUUM`, `ANALYZE`, and a few DDL commands such as `CREATE INDEX`"; `reltuples` carries the same estimate/update wording and "If the table has never yet been vacuumed or analyzed, `reltuples` contains `-1` indicating that the row count is unknown" +- https://www.postgresql.org/docs/current/functions-admin.html — `pg_relation_size(relation regclass [, fork text]) → bigint` "Computes the disk space used by one 'fork' of the specified relation ... With one argument, this returns the size of the main data fork"; results "are measured in bytes"; `pg_table_size` "Computes the disk space used by the specified table, excluding indexes (but including its TOAST table if any, free space map, and visibility map)" +- https://www.postgresql.org/docs/current/runtime-config-preset.html — `block_size`: "Reports the size of a disk block. It is determined by the value of `BLCKSZ` when building the server. The default value is 8192 bytes" +- https://www.postgresql.org/docs/current/sql-analyze.html — "For large tables, `ANALYZE` takes a random sample of the table contents, rather than examining every row"; "the statistics are only approximate, and will change slightly each time `ANALYZE` is run"; the collected statistics are "a list of some of the most common values in each column and a histogram showing the approximate data distribution" +- https://www.postgresql.org/docs/current/monitoring-stats.html — `pg_stat_all_tables`: `n_live_tup` "Estimated number of live rows", `n_mod_since_analyze` "Estimated number of rows modified since this table was last analyzed", `last_analyze` "Last time at which this table was manually analyzed", `last_autoanalyze` "Last time at which this table was analyzed by the autovacuum daemon" +- https://www.postgresql.org/docs/current/ddl-system-columns.html — `ctid`: "The physical location of the row version within its table. Note that although the `ctid` can be used to locate the row version very quickly, a row's `ctid` will change if it is updated or moved by `VACUUM FULL`" +- https://www.postgresql.org/docs/current/datatype-oid.html — "`tid`, or tuple identifier (row identifier) ... A tuple ID is a pair (block number, tuple index within block) that identifies the physical location of the row within its table" +- https://www.postgresql.org/docs/release/14.0/ — Optimizer: "Allow efficient heap scanning of a range of `TIDs` (Edmund Horner, David Rowley) ... Previously a sequential scan was required for non-equality `TID` specifications" +- Field observation 2026-08-12 (PostgreSQL, `external_data.collected_from_seumter`, ~1.07M rows, no index on the batch-date column): `relpages` = 43992 while `pg_relation_size` gave 44897 blocks — 905 blocks appended since the last maintenance. `ctid > '(43970,0)' LIMIT 20` returned only `20260619`, while `SELECT DISTINCT` over `ctid > '(44880,0)'` returned `20260719`. The `pg_stats` MCV list topped out at `20260427`, missing three months of arrivals; `reltuples` (1,069,782) trailed `n_live_tup` (1,092,465) by ~22.7k rows, which was the cross-signal that the catalog was behind the heap diff --git a/wiki/databases/data-survey/surveying-live-data-for-a-rule.md b/wiki/databases/data-survey/surveying-live-data-for-a-rule.md index 604279f..49baaa6 100644 --- a/wiki/databases/data-survey/surveying-live-data-for-a-rule.md +++ b/wiki/databases/data-survey/surveying-live-data-for-a-rule.md @@ -8,7 +8,7 @@ sources: - https://www.postgresql.org/docs/current/functions-aggregate.html - https://greatexpectations.io/blog/exploring-data-quality-volume/ last_verified: 2026-08-05 -related: [databases-query-optimization-existence-and-count-checks, databases-schema-design-requirements-to-tables, databases-schema-design-nullability-and-defaults] +related: [databases-data-survey-catalog-statistics-as-current-state, databases-query-optimization-existence-and-count-checks, databases-schema-design-requirements-to-tables, databases-schema-design-nullability-and-defaults] --- # Surveying Live Data to Derive a Mapping or Normalization Rule diff --git a/wiki/databases/index.md b/wiki/databases/index.md index fc3e9e9..59e769a 100644 --- a/wiki/databases/index.md +++ b/wiki/databases/index.md @@ -14,6 +14,7 @@ Match your situation to a "load when" line; load only matching pages. | [composite-index-column-order](indexing/composite-index-column-order.md) | Creating a multi-column index; choosing column order for equality + range/sort queries | | [covering-indexes](indexing/covering-indexes.md) | A query already served by an index still reads the table (heap) heavily; deciding whether to add INCLUDE/covering columns | | [partial-and-expression-indexes](indexing/partial-and-expression-indexes.md) | Queries always filter a fixed rare condition (status, deleted_at) or a function of a column (lower(email)); a uniqueness rule applies only to a subset of rows (e.g. live rows) | +| [trigram-index-short-patterns](indexing/trigram-index-short-patterns.md) | A `LIKE`/`ILIKE '%keyword%'` search on a PostgreSQL `pg_trgm` GIN/GiST index is fast for ordinary words and slow for one- or two-character keywords; `EXPLAIN` shows a `Bitmap Index Scan` on the trigram index and the query is still slow; choosing a minimum search-keyword length, or deciding between pg_trgm, pg_bigm, and a driver index for another condition | | [index-write-cost](indexing/index-write-cost.md) | Adding indexes to write-heavy tables; bulk loads; auditing for unused/redundant indexes | ## query-optimization @@ -21,6 +22,7 @@ Match your situation to a "load when" line; load only matching pages. | Page | Load when | |------|-----------| | [reading-execution-plans](query-optimization/reading-execution-plans.md) | A single query/statement is slow; verifying an index/query change with EXPLAIN before shipping (endpoint slow because it runs *many* fast queries → n-plus-one-queries) | +| [comparing-two-execution-plans](query-optimization/comparing-two-execution-plans.md) | Attributing a slowdown to one variable by comparing `EXPLAIN (ANALYZE)` across two variants of a statement; one arm came back far faster or its plan shows `never executed` / `loops=0`; quoting the duration of an arm a client deadline cut short | | [keyset-pagination](query-optimization/keyset-pagination.md) | Implementing pagination, infinite scroll, or batch table walks | | [streaming-large-result-sets](query-optimization/streaming-large-result-sets.md) | Exporting/reading a very large single-query result into the app; process memory peaks on `fetchall` or building a big file; server-side cursor blocked by autocommit or a read-only proxy | | [large-in-lists](query-optimization/large-in-lists.md) | Building `IN (...)` queries whose list size can grow (batch lookups, fetch-by-ids) | @@ -52,6 +54,8 @@ Match your situation to a "load when" line; load only matching pages. | Page | Load when | |------|-----------| +| [catalog-statistics-as-current-state](data-survey/catalog-statistics-as-current-state.md) | Finding a large PostgreSQL table's newest rows when the column has no index, so you reach for `pg_class.relpages`/`reltuples`, `pg_stats` most-common values, or a `ctid` range scan of the last blocks; about to conclude "no recent data" from a cheap catalog probe; choosing the scan's starting block or how to bound a `ctid` range | +| [audit-columns-as-update-evidence](data-survey/audit-columns-as-update-evidence.md) | About to read `update_dt`/`updated_at`/`modified_by` (NULL, or equal to the insert timestamp) as evidence that rows were never modified or a feature is unused; an incident needs to know whether writes to those rows were *attempted*; deciding whether ORM callback auditing or a DB trigger is the right writer for the claim you need to make | | [surveying-live-data-for-a-rule](data-survey/surveying-live-data-for-a-rule.md) | A task says to sample real data to decide a mapping table, normalization/canonicalization rule, enum value set, or parsing rule; a `GROUP BY`/`DISTINCT` survey came back with zero rows; deciding what evidence replaces the data when the table is empty; recording in the deliverable which evidence a rule was actually derived from | ## sqlite diff --git a/wiki/databases/indexing/index-selection.md b/wiki/databases/indexing/index-selection.md index f5a92c8..f27fe56 100644 --- a/wiki/databases/indexing/index-selection.md +++ b/wiki/databases/indexing/index-selection.md @@ -8,7 +8,7 @@ sources: - https://www.postgresql.org/docs/current/indexes.html - https://use-the-index-luke.com/ last_verified: 2026-07-10 -related: [databases-indexing-composite-index-column-order, databases-indexing-index-write-cost, databases-query-optimization-reading-execution-plans] +related: [databases-indexing-composite-index-column-order, databases-indexing-index-write-cost, databases-query-optimization-reading-execution-plans, databases-indexing-trigram-index-short-patterns] --- # Deciding Whether a Column Needs an Index @@ -48,7 +48,7 @@ designing a new table and choosing initial indexes. |------|------| | Table is small (fits in a few pages) | Planner will sequential-scan regardless; skip the index until the table grows | | Column has few distinct values but you always query one rare value | Partial index on that value beats a full index | -| Text search / `LIKE '%term%'` | B-tree cannot serve infix matches; use a trigram or full-text index type instead of adding a useless B-tree | +| Text search / `LIKE '%term%'` | B-tree cannot serve infix matches; use a trigram or full-text index type instead of adding a useless B-tree — and set the minimum keyword length that index type needs ([databases-indexing-trigram-index-short-patterns]) | | Write-heavy table, marginal read gain | Weigh maintenance cost first ([databases-indexing-index-write-cost]) | | Creating the index on a large live table | PostgreSQL: `CREATE INDEX CONCURRENTLY` — no long write-block, cannot run inside a transaction, and a failed build leaves an `INVALID` index (drop it, retry). MySQL 8.0: online DDL (`ALGORITHM=INPLACE, LOCK=NONE`) | diff --git a/wiki/databases/indexing/trigram-index-short-patterns.md b/wiki/databases/indexing/trigram-index-short-patterns.md new file mode 100644 index 0000000..1dd0581 --- /dev/null +++ b/wiki/databases/indexing/trigram-index-short-patterns.md @@ -0,0 +1,95 @@ +--- +id: databases-indexing-trigram-index-short-patterns +domain: databases +category: indexing +applies_to: [postgresql] +confidence: verified +sources: + - https://www.postgresql.org/docs/current/pgtrgm.html + - https://postgrespro.com/list/thread-id/1821635 + - https://github.com/pgbigm/pg_bigm/blob/master/docs/pg_bigm_en.md +last_verified: 2026-08-10 +related: + [ + databases-indexing-index-selection, + databases-query-optimization-reading-execution-plans, + databases-indexing-partial-and-expression-indexes, + databases-indexing-covering-indexes, + ] +--- + +# Substring Search on a Trigram Index When the Keyword Is Shorter Than Three Characters + +## When this applies + +A `LIKE`/`ILIKE '%keyword%'` search is served by a PostgreSQL `pg_trgm` GIN or +GiST index, and the keyword comes from a user — so it can be one or two +characters. Also when such a search is fast for ordinary words and slow for short +ones, or when `EXPLAIN` shows a `Bitmap Index Scan` on the trigram index and the +query is still slow, or when you are choosing the minimum length for a search +input. + +Reading the plan that shows this → [databases-query-optimization-reading-execution-plans]. +Choosing the index type in the first place → [databases-indexing-index-selection]. + +## Do this + +1. **Count the characters between wildcards, not the characters the user typed.** + The index is searched by extracting trigrams from the pattern, and "a pattern + with no extractable trigrams will degenerate to a full-index scan". A + wildcard-delimited segment of fewer than three characters yields none: + `get_wildcard_trigrams` "return[s] no trigrams for wildcard part 'st' since + charlen < 3", so "GIN_SEARCH_MODE_ALL mode is used and results in full index + scan instead of trigrams being used". `show_trgm('cat')` returning four + trigrams does not contradict this — that padding applies to a *word* being + indexed, and a `%…%` pattern asserts no word boundary to pad against. + +2. **Read the cost from the recheck counters, not from the scan node's name.** + The plan still reads `Bitmap Index Scan` on the trigram index; what changes is + that the candidate bitmap becomes every row, so the work moves into the heap + recheck. `EXPLAIN (ANALYZE, BUFFERS)` is what shows it — compare + `Rows Removed by Index Recheck` against the table's row count and the buffer + count against the short and long pattern. + +3. **Pick the branch by whether short keywords are a supported input:** + +| Situation | Do | +|-----------|-----| +| The search field has no other selective filter and short keywords are optional | Enforce a minimum keyword length at the API boundary and return a stated validation error, so the cost is refused rather than paid | +| The same query carries another selective condition (owner, department, tenant, date range) | Give that condition its own index and let it produce the bitmap, then let the substring match run as a heap filter — this bounds the scan by the selective condition instead of the pattern | +| Short keywords must return results and the operator is `LIKE` | Evaluate `pg_bigm`, which "allows a user to create **2-gram** (bigram) index", and whose own comparison rates "Full text search with 1-2 characters keyword" as "Fast" against pg_trgm's "Slow" — footnoted with the same mechanism, "only sequential scan or index full scan (not normal index scan) can run" | +| Short keywords must return results and the query needs `ILIKE`, `~`, or `~*` | Keep pg_trgm and normalize instead — index and query one case-folded expression ([databases-indexing-partial-and-expression-indexes]) — because pg_bigm's index supports "LIKE only" while pg_trgm supports "LIKE (~~), ILIKE (~~*), ~, ~*" | +| The short keyword is a prefix, not an infix (`'ab%'`) | Serve it from a B-tree on the column (or its case-folded expression) — a left-anchored pattern needs no trigrams | + +4. **Verify the chosen branch on production-scale data before shipping it.** The + degeneration is invisible at small row counts, where a full index scan is + cheap; measure at the table's real size. + +## Edge cases + +| Case | Then | +|------|------| +| The other WHERE conditions appear under `Bitmap Heap Scan` as `Filter` rather than as `Index Cond` | They are not reducing the scan — they are applied after the rows are read, so the plan is still paying the full recheck. Add the index that lets one of them drive the bitmap | +| The query returns very few rows, so the result looks cheap | Read the cost from buffers and recheck counts, not from the row count — the scan reads the whole index and rechecks the whole heap whichever way the match comes out | +| Only some of the keyword's wildcard segments are short (`'%ab%defg%'`) | The pattern has extractable trigrams from the longer segment, so the index search works; the short segment contributes nothing and is checked on recheck | +| The column is searched with both a short and a long keyword in one `OR` | The short branch degenerates independently; split the branches so the long one keeps its index path, or apply the length rule per branch | +| The table is small today and the search is new | Record the row count at which the branch was chosen — the same query flips from acceptable to a full-table recheck with growth, and nothing in the plan's shape changes when it does | +| A GiST trigram index is used instead of GIN | The same extraction rule governs it: with no extractable trigrams there is nothing to look up, and the docs' degeneration statement covers "both `LIKE` and regular-expression searches" | +| The workload is non-alphabetic text (Japanese, Chinese, Korean) | The same comparison lists pg_trgm's full text search for non-alphabetic language as "Not supported", so the 3-character rule bites ordinary two-character words — treat pg_bigm as the default candidate rather than the fallback. Its footnote records the alternative, "commenting out KEEPONLYALNUM macro variable in contrib/pg_trgm/pg_trgm.h and rebuilding pg_trgm module", which makes the choice a build-vs-extension decision rather than a capability wall | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Read `Bitmap Index Scan` on the trigram index as proof the index is doing the work | Compare `Rows Removed by Index Recheck` with the table's row count | The node name is the same in both cases; the recheck count is what separates a lookup from a full scan | +| Add a second trigram index, or `REINDEX`, because the short-keyword query is slow | Apply the length rule from step 3 | The index is being scanned in full by design for a pattern with no trigrams; another copy of it is scanned in full too | +| Raise `work_mem` or add heap-side tuning to make the short query fit | Refuse the short pattern at the boundary or give the query a selective driver index | The cost is proportional to the table, not to the memory available for the bitmap | +| Assume a two-character search is cheap because it returns three rows | Measure buffers for a two-character and a three-character pattern on the same index | Measured 2026-08-10 on a 4.64M-row table: a 3-character `ILIKE` ran 17 ms / 4 buffers, the 2-character one 18,789 ms / 121,837 buffers, with `Rows Removed by Index Recheck: 4,640,486` and 3 rows matched | +| Swap pg_trgm for pg_bigm to fix a slow `ILIKE` | Decide the case-folding strategy first, then choose | pg_bigm's operator support is "LIKE only"; an `ILIKE` workload has to be rewritten to a normalized expression either way | + +## Sources + +- https://www.postgresql.org/docs/current/pgtrgm.html — "For both `LIKE` and regular-expression searches, keep in mind that a pattern with no extractable trigrams will degenerate to a full-index scan"; "The index search works by extracting trigrams from the search string and then looking these up in the index. The more trigrams in the search string, the more effective the index search is"; "A trigram is a group of three consecutive characters taken from a string"; and the padding rule — "Each word is considered to have two spaces prefixed and one space suffixed when determining the set of trigrams contained in the string" — which is why `show_trgm` on a short *word* still returns trigrams while a `%…%` pattern yields none +- https://postgrespro.com/list/thread-id/1821635 — Amit Langote, pgsql list thread (2013-05-31): "When I debugged a partial match case such as 'column like '%st%'', it appears that get_wildcard_trigrams return no trigrams for wildcard part 'st' since charlen < 3"; "Hence, GIN_SEARCH_MODE_ALL mode is used and results in full index scan instead of trigrams being used". This is the mechanism behind the docs' one-sentence statement, and it is stated in terms of the wildcard-delimited segment rather than the whole pattern +- https://github.com/pgbigm/pg_bigm/blob/master/docs/pg_bigm_en.md — "The pg_bigm module provides full text search capability in [PostgreSQL]. This module allows a user to create **2-gram** (bigram) index for faster full text search." Its pg_trgm comparison table (verified against the raw file 2026-08-10, cell by cell) reads: "Phrase matching method for full text search" 3-gram vs 2-gram; "Available text search operators" "LIKE (~~), ILIKE (~~*), ~, ~*" vs "LIKE only"; "Full text search for non-alphabetic language (e.g., Japanese)" "Not supported (\*1)" vs "Supported"; "Full text search with 1-2 characters keyword" "Slow (\*2)" vs "Fast"; "Available index" "GIN and GiST" vs "GIN only". Footnote (\*2) gives the mechanism independently of the PostgreSQL docs — "Because, in this search, only sequential scan or index full scan (not normal index scan) can run" — and footnote (\*1) records that pg_trgm's non-alphabetic limit is liftable "by commenting out KEEPONLYALNUM macro variable … and rebuilding pg_trgm module". The operator row is the constraint that decides step 3's last two rows +- Field measurement 2026-08-10 (PostgreSQL, 4,640,489-row table, `gin(tip_ctn gin_trgm_ops)`, `EXPLAIN (ANALYZE, BUFFERS)`): a 3-character `ILIKE '%…%'` ran 17 ms reading 4 buffers; a 2-character `ILIKE '%TI%'` on the same index and column ran 18,789 ms reading 121,837 buffers with `Rows Removed by Index Recheck: 4,640,486` and 3 rows actually matching. Both plans showed a `Bitmap Index Scan` on the trigram index, and the query's other conditions appeared as `Filter` on the `Bitmap Heap Scan`, reducing nothing diff --git a/wiki/databases/query-optimization/comparing-two-execution-plans.md b/wiki/databases/query-optimization/comparing-two-execution-plans.md new file mode 100644 index 0000000..8e2c57e --- /dev/null +++ b/wiki/databases/query-optimization/comparing-two-execution-plans.md @@ -0,0 +1,109 @@ +--- +id: databases-query-optimization-comparing-two-execution-plans +domain: databases +category: query-optimization +applies_to: [postgresql] +confidence: verified +sources: + - https://www.postgresql.org/docs/current/using-explain.html + - https://www.postgresql.org/docs/current/ddl-partitioning.html + - https://github.com/postgres/postgres/blob/083ac033419f690758508e08c1736089384bbee8/src/backend/commands/explain.c + - https://www.postgresql.org/docs/current/pgstatstatements.html + - https://www.postgresql.org/docs/current/monitoring-stats.html +last_verified: 2026-08-11 +related: + [ + databases-query-optimization-reading-execution-plans, + databases-indexing-index-selection, + debugging-methodology-hypothesis-testing, + ] +--- + +# Attributing a Slowdown to One Variable Across Two Execution Plans + +## When this applies + +You ran `EXPLAIN (ANALYZE)` on two variants of the same statement — literal vs +bound parameter, index on vs off, filter A vs filter B — and are about to say +"the slowness is caused by X" because one arm was far faster, or to publish an +arm's duration when a client-side deadline cut that arm short. + +Reading a single plan → [databases-query-optimization-reading-execution-plans]. + +## Do this + +1. **Read the fast arm's plan for `never executed` before attributing anything.** + PostgreSQL prints ` (never executed)` whenever a node has no instrumentation + with a positive loop count — under `EXPLAIN (ANALYZE)` on an ordinary plan + node that means it ran zero times. An arm whose expensive subtree never + executed did not pay the cost you are comparing against; it is a different + experiment, not a baseline. +2. **Count the variables that actually moved.** When the arms differ in both + "the suspected factor X" and "whether rows reached the expensive subtree", + two hypotheses explain the gap identically — X, and plain row volume. Any + attribution to X alone is unsupported at that point. +3. **Add a third arm that holds X at the fast arm's setting and makes rows + flow.** Only the pair that differs in X *with rows flowing in both* attributes + the difference to X. Record every arm with the **actual row count** its plan + reports, not a yes/no: the attribution holds only when the two compared arms + moved comparable row volumes, and a bare "yes" hides an arm that passed ten + rows against another that passed a million. + +| Arm | X | Rows reach the expensive subtree | What it establishes | +|-----|---|----------------------------------|---------------------| +| 1 | fast setting | no | Confounded — reports the cost of skipping, not of X | +| 2 | fast setting | yes | The baseline arm 1 was mistaken for | +| 3 | suspected setting | yes | Compared against arm 2, isolates X | +| 4 | suspected setting | no | Separates "X alone" from "X plus a specific subtree" | + +4. **Quote only durations that came from a completed execution.** A ">25 s" + from a client that gave up is a property of the client's deadline, not of the + query. Recover the real number by re-running the arm to completion with the + client deadline removed, and read `now() - query_start` from + `pg_stat_activity` for that backend while it runs if you need the number + before it finishes (`query_start` is "Time when the currently active query was + started"). +5. **Check the arms are otherwise equal** — same data, same instance, and the + caches in the same state — before scoring the gap + ([databases-query-optimization-reading-execution-plans] covers the warm-cache + trap). + +## Edge cases + +| Case | Then | +|------|------| +| You are reading `EXPLAIN (ANALYZE, FORMAT JSON)` or feeding plans to a script | The string `never executed` exists only in TEXT format; JSON/XML/YAML emit `"Actual Loops": 0` instead. Gate the check on `Actual Loops == 0` — that field is emitted unconditionally, while `Actual Total Time` appears only when timing is on, so a parser keyed on the time silently scores a skipped subtree as a 0 ms one | +| You want the cancelled arm's duration from `pg_stat_statements` | It is not there. The view accumulates execution statistics "only for successful operations", so a statement cancelled by the client or by `statement_timeout` contributes nothing to `calls`/`total_exec_time` — a number you do find for that query text came from some *other*, completed run | +| You reach for `auto_explain` instead, to capture the cancelled arm | It has the same blind spot for the same reason: `auto_explain` logs from `explain_ExecutorEnd`, which it installs as `ExecutorEnd_hook`, and a cancelled statement raises `ERROR` before reaching `ExecutorEnd`. Point `auto_explain` at the deadline-free **re-run** instead of at the cancelled arm | +| A duration you did find in `pg_stat_statements` looks plausible | Divide `total_exec_time` by `calls` before comparing; the column is a running total across every completed execution. It is named `total_time` in extension version 1.7 and earlier — which ships with PostgreSQL 12 and earlier, and also persists on a newer server whose extension was never `ALTER EXTENSION pg_stat_statements UPDATE`d | +| The plan node count differs between arms, not just the timings | The planner chose different shapes; compare the per-node actual times rather than the totals, and treat the shape change itself as the finding | +| Only one arm can be run against production | Run the confounded-arm check anyway — `never executed` is visible in the single plan you have, and it tells you the measurement is not a cost | +| The fast arm's subtree is skipped because a filter genuinely matches nothing in production too | That is a real optimization, not a confound — state it as "fast when the filter is empty", and keep arm 2 to document the non-empty cost | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Conclude "X is the cause" from a two-arm gap of 37 ms vs 143 s | Add the third arm with rows flowing and re-compare | The 37 ms arm may never have run the pipeline; the gap then measures skipping, not X | +| Read a fast plan's small total time as "this plan is efficient" | Scan for `never executed` / `loops=0` first | Zero executions is the cheapest possible plan and tells you nothing about the plan's cost | +| Publish ">25 s" for an arm the client cancelled | Re-run it to completion without the deadline and publish the measured value | A client timeout bounds the number from above; quoting it understates the cost by an unknown amount | +| Fix the filter so rows start flowing, having only ever measured the zero-row case | Measure the rows-flowing cost first | The fix changes which arm production runs; nobody has priced the arm you are about to ship | +| Treat "one variable per experiment" as satisfied because you edited one token | Verify in the plan that only one variable moved | The planner can change a second variable — whether a subtree runs at all — in response to your one edit ([debugging-methodology-hypothesis-testing]) | + +## Sources + +- https://www.postgresql.org/docs/current/using-explain.html — `EXPLAIN ANALYZE` reports actual row counts and loops per node; plain `EXPLAIN` shows intent only +- https://www.postgresql.org/docs/current/ddl-partitioning.html — the one official page that names the marker: "Determining if partitions were pruned during this phase requires careful inspection of the `loops` property in the `EXPLAIN ANALYZE` output. … Some may be shown as `(never executed)` if they were pruned every time." (`using-explain.html` does not mention it) +- https://github.com/postgres/postgres/blob/083ac033419f690758508e08c1736089384bbee8/src/backend/commands/explain.c — `ExplainNode` prints `" (actual time=… rows=… loops=…)"` only under `if (es->analyze && planstate->instrument && planstate->instrument->nloops > 0)`; the `else if (es->analyze)` branch appends `" (never executed)"` in `EXPLAIN_FORMAT_TEXT`, and in every other format emits `Actual Rows` and `Actual Loops` of `0` unconditionally plus `Actual Startup Time`/`Actual Total Time` of `0.0` when `es->timing` is set — so non-text formats carry no such string, and `Actual Loops` is the field always present to test (read at commit `083ac03`, the tip of branch `REL_17_STABLE` at the time — `REL_17_STABLE` is a branch, not a tag, so the URL is pinned to the SHA; lines 1841–1888) +- https://www.postgresql.org/docs/current/pgstatstatements.html — "planning and execution statistics are updated at their respective end phase, and only for successful operations"; the execution columns are `calls`, `total_exec_time`, `mean_exec_time` +- https://www.postgresql.org/docs/current/monitoring-stats.html — `pg_stat_activity.query_start` is "Time when the currently active query was started, or if `state` is not `active`, when the last query was started" +- Field measurement 2026-08-11 (one statement, same data and instance, four arms; X = whether the filter value reached the planner as a literal or as an opaque bound parameter): + +| Arm | X | Rows reach the aggregate subqueries | Duration | +|-----|---|-------------------------------------|----------| +| 1 | literal | no — every aggregate subquery printed `never executed` | 37.7 ms | +| 2 | literal | yes | 7,960 ms | +| 3 | opaque | yes | 143,658 ms | +| 4 | opaque | no | 1,859 ms | + + This record keeps only the binary "did rows reach the subqueries", not the per-arm row counts directive 3 asks for — so it supports the qualitative attribution below but not a claim that arms 2 and 3 moved equal volumes. The two-arm reading available at the time was arm 1 vs arm 3 (37.7 ms vs 143,658 ms), which attributed the whole gap to parameter opacity. Arm 4 falsified that: opacity with nothing flowing costs 1,859 ms. The attributable comparison is arm 2 vs arm 3 — same rows flowing, X the only difference — an 18× penalty that appears only once the aggregate subqueries run. Arm 3's duration was first known only as ">25 s" because the client cancelled; it was recovered by re-running the arm without the client deadline, not from `pg_stat_statements`, which held no row for the cancelled execution diff --git a/wiki/databases/query-optimization/reading-execution-plans.md b/wiki/databases/query-optimization/reading-execution-plans.md index 2c6a90c..72395a1 100644 --- a/wiki/databases/query-optimization/reading-execution-plans.md +++ b/wiki/databases/query-optimization/reading-execution-plans.md @@ -8,7 +8,7 @@ sources: - https://www.postgresql.org/docs/current/using-explain.html - https://dev.mysql.com/doc/refman/8.0/en/explain-output.html last_verified: 2026-07-10 -related: [databases-indexing-index-selection, databases-query-optimization-large-in-lists] +related: [databases-indexing-index-selection, databases-query-optimization-large-in-lists, databases-query-optimization-comparing-two-execution-plans] --- # Diagnosing a Query with Its Execution Plan @@ -45,6 +45,7 @@ before shipping it. | Plan differs between parameter values | Skewed data: one plan per value class. Test with production-representative parameters, worst class included | | `EXPLAIN ANALYZE` on writes (`INSERT/UPDATE/DELETE`) | It executes them. Wrap in `BEGIN; ... ROLLBACK;` | | Production incident, can't run ANALYZE variants freely | Capture the live plan via `pg_stat_statements` / slow query log + `auto_explain` instead of experimenting on prod | +| A node prints `never executed`, or you are comparing this plan against a second variant to blame one variable | The node ran zero times, so its cost is unmeasured — see [databases-query-optimization-comparing-two-execution-plans] before attributing the difference | ## Sources diff --git a/wiki/debugging/index.md b/wiki/debugging/index.md index 8452e6f..521e247 100644 --- a/wiki/debugging/index.md +++ b/wiki/debugging/index.md @@ -12,6 +12,7 @@ Match your situation to a "load when" line; load only matching pages. | [reproduce-first](methodology/reproduce-first.md) | A bug is reported or behavior is wrong and you are about to investigate or fix; deciding what to capture when full reproduction is impossible (prod-only, timing-dependent, one-off crash) | | [isolate-by-bisection](methodology/isolate-by-bisection.md) | A bug reproduces but its location is unknown; it worked before / works in env A but not env B / fails with one input but not another — binary-searching versions (git bisect), code paths, data, or environment diffs | | [hypothesis-testing](methodology/hypothesis-testing.md) | You have a suspect cause and are about to "try a fix"; several suspects compete and you must pick what to test next; verifying that a fix that "worked" actually addressed the mechanism | +| [probe-path-vs-operation-path](methodology/probe-path-vs-operation-path.md) | A precondition probe (login status, health, connectivity) reports success while the operation it gates fails with an auth/permission error; a browser page-load login check gates direct API calls made with stored cookies; deciding what a preflight probe must exercise under refresh-token cookie auth | | [verify-the-fix](methodology/verify-the-fix.md) | You believe a bug is fixed and are about to close or ship it; the bug "cannot be reproduced anymore" after changes; a previously fixed bug came back; deciding what must pass (repro re-run, both directions, regression test) and what to clean up before closing | ## signals diff --git a/wiki/debugging/methodology/hypothesis-testing.md b/wiki/debugging/methodology/hypothesis-testing.md index 282859c..818f770 100644 --- a/wiki/debugging/methodology/hypothesis-testing.md +++ b/wiki/debugging/methodology/hypothesis-testing.md @@ -7,8 +7,8 @@ confidence: verified sources: - https://www.debuggingbook.org/html/Intro_Debugging.html - https://sre.google/sre-book/effective-troubleshooting/ -last_verified: 2026-07-10 -related: [debugging-methodology-reproduce-first, debugging-methodology-isolate-by-bisection] +last_verified: 2026-08-06 +related: [debugging-methodology-reproduce-first, debugging-methodology-isolate-by-bisection, qa-deliverables-exclusivity-and-absence-claims, debugging-methodology-probe-path-vs-operation-path] --- # Testing a Suspected Cause Before Changing Code @@ -53,6 +53,8 @@ applies when several suspects compete and you must pick what to investigate next | Every hypothesis you can think of is falsified | Your model of the system is wrong somewhere upstream. Return to evidence gathering: widen what you observe ([debugging-signals-logs-and-correlation]) or bisect to relocate the fault ([debugging-methodology-isolate-by-bisection]) | | Testing the hypothesis requires touching prod | Prefer a read-only prediction (something already in logs/metrics that must be true if the hypothesis holds); active prod experiments require owner approval and a rollback plan | | Hypothesis confirmed but the fix belongs to another domain (slow query, infra limit) | Record the confirmed mechanism, then route the fix to the owning domain — e.g. a slow SQL statement goes to wiki/databases/query-optimization/reading-execution-plans.md | +| One suspect is a mechanism designed to be invisible to you (silent moderation, shadowban, silent spam-drop, a filter that returns success) | An explicit error message refutes that suspect: concealment is the mechanism's defining property, so a system that told you it refused you is a different system. Test the visible-refusal suspects first — they are the ones with readable evidence | +| Your only evidence is a query that returned zero results | Zero is consistent with "the thing was removed" and with "the thing was never accepted", so it cannot separate them. Run a positive control through the same query path — an input you know must return rows — to establish the path works, then find evidence that differs between the two suspects | ## Instead of @@ -67,3 +69,5 @@ applies when several suspects compete and you must pick what to investigate next - https://www.debuggingbook.org/html/Intro_Debugging.html — scientific method for debugging: hypothesis → prediction → experiment → repeat - https://sre.google/sre-book/effective-troubleshooting/ — hypothetico-deductive troubleshooting; test causes, treat the confirmed one +- https://news.ycombinator.com/newsfaq.html — a killed post is marked `[dead]` and "aren't displayed by default"; the FAQ documents no notification to the author, i.e. the suppression carries no message to the affected party — which is why receiving an explicit refusal falsifies the silent-suppression hypothesis +- Field application 2026-08-06: a submission was assumed silently suppressed; the account had in fact received an explicit restriction notice, and a search API returning zero hits for the item was shown to be uninformative by a control query on the same endpoint returning 8,462 hits — the control separated "endpoint broken" from "item absent" diff --git a/wiki/debugging/methodology/probe-path-vs-operation-path.md b/wiki/debugging/methodology/probe-path-vs-operation-path.md new file mode 100644 index 0000000..fcb201c --- /dev/null +++ b/wiki/debugging/methodology/probe-path-vs-operation-path.md @@ -0,0 +1,67 @@ +--- +id: debugging-methodology-probe-path-vs-operation-path +domain: debugging +category: methodology +applies_to: [general, headless-browser, cookie-auth] +confidence: verified +sources: + - https://github.com/velopert/velog-server/blob/master/src/lib/token.ts +last_verified: 2026-08-14 +related: [debugging-methodology-hypothesis-testing, infrastructure-agent-orchestration-control-signals-vs-primary-artifacts] +--- + +# A Passing Precondition Probe for a Failing Operation + +## When this applies + +An automation's precondition check (login status, health, connectivity) reports +success, yet the operation it gates fails immediately after with an +auth/permission error — e.g. a headless-browser "logged in" check passes while +the API call made with the same stored cookies returns "Not logged in". Also +applies when deciding what a preflight probe for a direct API operation must +exercise. + +## Do this + +1. **Treat a probe's success as evidence about the probe's path only.** Before + trusting it, state what the probe actually exercised and diff that against + the failing operation: same credential material, same endpoint/host, same + client. A green probe over a different path does not contradict the failure + — it locates it. +2. **Probe the operation's own path with the operation's exact inputs.** For a + cookie/token-authenticated API, call the API's identity endpoint (GraphQL + `currentUser`, REST `/me`) with the same cookie jar and client the operation + will use, and require a non-null identity before proceeding. +3. **Under refresh-token cookie auth, expect page-load probes and replayed + cookie jars to diverge.** The server keeps a short-lived access token beside + a long-lived refresh token and rotates both via `Set-Cookie` when a request + arrives with an expired access token. A browser context persists that + rotation, so a page-load login check passes for as long as the refresh token + lives; a script replaying a stored cookie jar keeps sending the stale access + token it saved. +4. **On an API-probe failure, refresh the stored credentials, persist them, and + re-probe** — re-run the login flow (or the refresh endpoint), write the + rotated cookies back to the store the operation reads, then repeat step 2 + before the operation. + +## Edge cases + +| Case | Then | +|------|------| +| Probe page and operation API live on different hosts (`velog.io` page vs `v3.velog.io/graphql`) | Cookie domain scoping can differ per host — verify the jar's cookies actually attach to the API request (dump request headers), not just that they exist in the store | +| The API layer itself refreshes when handed a valid refresh token (velog's `consumeUser` middleware) | Rotation is returned via `Set-Cookie`; a client that discards response cookies works once and fails on a later run — persist rotated cookies after every authenticated call | +| Probe passes, operation starts, then fails auth mid-run | The access token expired during the operation; capture the operation's own error and re-authenticate there — tightening the preflight cannot cover a token that outlives it | +| Both probe and operation fail after refresh | The refresh token itself is expired or revoked — re-run the interactive login flow, not the refresh path | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Gate a direct API operation on a browser page-load login check | Gate on the API's identity query with the operation's own client and cookie jar | Page navigation triggers server-side token refresh that the browser persists; the check validates the refresh token while the operation depends on the stored access token | +| Retry the operation because the probe "proves" auth is fine | Diff the probe's path against the operation's path first | The contradiction is the diagnosis: two paths, one credential expired on exactly one of them | +| Fix the failure by loosening or removing the probe | Move the probe onto the operation's path | The probe was not wrong, it was answering a different question | + +## Sources + +- https://github.com/velopert/velog-server/blob/master/src/lib/token.ts — `setTokenCookie` sets `access_token` with `maxAge` 1 hour beside `refresh_token` with 30 days; `consumeUser` middleware refreshes on an expired/near-expiry access token and returns new cookies via `Set-Cookie`, so only clients that persist response cookies stay authenticated +- Field reproduction 2026-08-14 (auto-velog pipeline, headless Chromium + stored cookie jar): page-load check returned `STATUS:LOGGED_IN`, the immediately following publish mutation failed with "Not logged in", and a direct `v3.velog.io/graphql` `currentUser` query with the same stored cookies returned null; re-running the login flow (rewriting the cookie store) made the same publish succeed diff --git a/wiki/debugging/methodology/reproduce-first.md b/wiki/debugging/methodology/reproduce-first.md index 36abd55..738bb84 100644 --- a/wiki/debugging/methodology/reproduce-first.md +++ b/wiki/debugging/methodology/reproduce-first.md @@ -8,8 +8,8 @@ sources: - https://sscce.org/ - https://www.debuggingbook.org/html/DeltaDebugger.html - https://sre.google/sre-book/effective-troubleshooting/ -last_verified: 2026-07-10 -related: [debugging-methodology-hypothesis-testing, debugging-concurrency-intermittent-failures] +last_verified: 2026-08-13 +related: [debugging-methodology-hypothesis-testing, debugging-concurrency-intermittent-failures, debugging-signals-logs-and-correlation] --- # Building a Reproduction Before Investigating a Bug @@ -54,6 +54,7 @@ When a full local reproduction is impossible, capture evidence instead: | Reproduction needs data you are not allowed to copy | Reproduce the shape, not the content: synthesize data matching the schema, volume, and the specific values named in the failure (nulls, empty lists, boundary sizes) | | The report names the exact line to fix | Reproduce anyway before editing; a reproduction that survives the claimed fix disproves the report's diagnosis cheaply | | Bug reproduces only on the reporter's machine | Diff the two environments one variable at a time — versions, locale, config — moving your environment toward theirs until it fails ([debugging-methodology-isolate-by-bisection]) | +| Prod-only failure with no visible error — the client swallows it (a `.catch()` that ignores, an empty error handler) and the action just "does nothing" | Grep the production service logs for the endpoint path before reading more code: from the UI a 500 and a no-op are indistinguishable, and one server-side exception line kills whole families of hypotheses that local code reading cannot ([debugging-signals-logs-and-correlation]) | ## Instead of @@ -68,3 +69,4 @@ When a full local reproduction is impossible, capture evidence instead: - https://sscce.org/ — minimal, self-contained example discipline - https://www.debuggingbook.org/html/DeltaDebugger.html — systematically reducing failure-inducing inputs - https://sre.google/sre-book/effective-troubleshooting/ — "simplify and reduce"; reproduction as the basis of diagnosis +- Field context 2026-08 (silent-swallow row, field-tested): a prod-only bookmark bug where backend code, proxy, and browser click were all verified normal from the outside; one `journalctl | grep bookmark` surfaced PostgreSQL's "no unique or exclusion constraint matching the ON CONFLICT specification", pinning the cause to a deployed DB left on an old schema — a cause invisible in the repo's code diff --git a/wiki/debugging/signals/reading-error-messages.md b/wiki/debugging/signals/reading-error-messages.md index 6dd651e..ac49f59 100644 --- a/wiki/debugging/signals/reading-error-messages.md +++ b/wiki/debugging/signals/reading-error-messages.md @@ -8,7 +8,7 @@ sources: - https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Errors - https://gcc.gnu.org/onlinedocs/gcc/Warning-Options.html last_verified: 2026-07-10 -related: [debugging-signals-stack-traces, debugging-methodology-hypothesis-testing] +related: [debugging-signals-stack-traces, debugging-methodology-hypothesis-testing, backend-common-errors-diagnostics-from-a-shared-code-path] --- # Reading an Error Message Before Acting on It diff --git a/wiki/frontend/data-fetching/async-ui-states.md b/wiki/frontend/data-fetching/async-ui-states.md index 636aaf1..9a10001 100644 --- a/wiki/frontend/data-fetching/async-ui-states.md +++ b/wiki/frontend/data-fetching/async-ui-states.md @@ -11,7 +11,7 @@ sources: - https://tanstack.com/query/latest/docs/framework/react/guides/optimistic-updates - https://react.dev/reference/react/Component last_verified: 2026-07-10 -related: [frontend-state-client-vs-server-state, frontend-data-fetching-race-conditions] +related: [frontend-state-client-vs-server-state, frontend-data-fetching-race-conditions, frontend-data-fetching-query-state-vs-fetch-state] --- # Designing Loading, Error, Empty, and Data States for an Async View @@ -59,6 +59,7 @@ Then apply these to the transitions between states: | List is empty because the user's filters excluded everything | Say so, and offer "clear filters" — the generic empty state ("add your first item") misleads | | Response resolves fast enough that the skeleton only flashes | Keep the reserved space but suppress indicator animation for sub-second responses — feedback that fast is distraction, not information | | Mutation has no inverse (send email, submit payment) | No optimistic update — render an explicit pending state until the server confirms | +| The query can be disabled (`enabled: false`) or paused offline, so it has no data and is not fetching | The four states above do not cover it — branch on the cache's status/fetchStatus pair ([frontend-data-fetching-query-state-vs-fetch-state]) | ## Instead of diff --git a/wiki/frontend/data-fetching/query-state-vs-fetch-state.md b/wiki/frontend/data-fetching/query-state-vs-fetch-state.md new file mode 100644 index 0000000..66ce45d --- /dev/null +++ b/wiki/frontend/data-fetching/query-state-vs-fetch-state.md @@ -0,0 +1,102 @@ +--- +id: frontend-data-fetching-query-state-vs-fetch-state +domain: frontend +category: data-fetching +applies_to: [react, tanstack-query] +confidence: verified +sources: + - https://tanstack.com/query/latest/docs/framework/react/guides/queries + - https://tanstack.com/query/latest/docs/framework/react/reference/useQuery + - https://tanstack.com/query/latest/docs/framework/react/guides/disabling-queries + - https://tanstack.com/query/latest/docs/framework/react/guides/network-mode +last_verified: 2026-08-07 +related: + [ + frontend-data-fetching-async-ui-states, + frontend-state-client-vs-server-state, + frontend-data-fetching-race-conditions, + testing-mocking-what-to-mock, + ] +--- + +# A Server-State Query That Has No Data and Is Not Loading + +## When this applies + +You are defining what a component receives from a TanStack Query hook and are +about to treat `data === undefined` as "loading". +Also when a view shows a permanent spinner with no error and no retry, or the +query it renders can be disabled (`enabled: false`, `skipToken`) or paused by +the network mode. + +Designing the loading / error / empty / data renderings themselves → +[frontend-data-fetching-async-ui-states]. + +## Do this + +1. **Take two independent inputs, not one.** TanStack Query exposes them as + separate fields for this reason: "The `status` gives information about the `data`: Do + we have any or not? The `fetchStatus` gives information about the `queryFn`: + Is it running or not?" — and "all combinations for `status` and `fetchStatus` + [are] possible". A component prop of `data | undefined` collapses both axes + into one bit and cannot recover them. + +2. **Branch on the combination, and give every cell a rendering:** + +| status | fetchStatus | What it means | Render | +| --------- | ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- | +| `pending` | `fetching` | First fetch in flight — this is `isLoading`, defined as "`isFetching && isPending`" | Skeleton | +| `pending` | `idle` | Disabled/lazy query: "status === 'pending' and fetchStatus === 'idle'" with `enabled: false` | The pre-request state — the prompt, the disabled form, the "select a row" placeholder | +| `pending` | `paused` | Wanted to fetch, has no connection: "state: 'pending', but fetchStatus: 'paused' if they are mounting for the first time, and you have no network connection" | Offline notice plus a retry affordance | +| `error` | `idle` | The attempt failed and nothing is retrying | Error message plus a retry affordance ([frontend-data-fetching-async-ui-states]) | +| `error` | `fetching` | A retry is already running over the failed state | Keep the error message, disable the retry affordance while it runs | +| `error` | `paused` | Failed, and the retry is waiting for a connection | Offline notice — a retry affordance here cannot run | +| `success` | `fetching` | Background refresh over existing data | The data, plus a subtle refresh indicator | +| `success` | `paused` | Data is on screen and a background refetch is waiting for a connection | The data, plus an offline/stale indicator instead of the refresh indicator | +| `success` | `idle` | Settled | The data (or the empty state when it is an empty collection) | + +3. **Pass the discriminator down.** Give the presentational component either the + two fields or an explicit union (`{kind: 'idle' | 'loading' | 'paused' | +'error' | 'ready', …}`) built at the boundary that holds the query. The union + makes every unhandled state a type error instead of a blank screen. + +4. **Use `isLoading` for spinners and `isPending` for "no data yet".** The docs + state the split directly: lazy queries "will be in `status: 'pending'` right + from the start because `pending` means that there is no data yet … you likely + cannot use this flag to show a loading spinner". + +5. **Cover the disabled and paused cells in tests explicitly.** A test that mocks + the query hook supplies the flags by hand and therefore only ever produces + combinations its author already thought of — so the combination that ships the + bug is the one the suite never constructs. Write one case per row of the step-2 + table, taking the flag values from that table rather than from the component + ([testing-mocking-what-to-mock]). + +## Edge cases + +| Case | Then | +| ---------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| The query has `initialData` or `placeholderData` | It starts at `status: 'success'`, so the pending rows never render — assert which data the user is seeing before treating success as authoritative | +| A disabled query has cached data from an earlier mount | It initializes as "status === 'success' or isSuccess" rather than pending; the idle-pre-request rendering does not apply | +| The component must trigger the fetch itself | Keep `enabled: false` rather than `skipToken`: "`refetch` from `useQuery` will not work with `skipToken`. Calling `refetch()` on a query that uses `skipToken` will result in a `Missing queryFn` error" | +| `select` narrows the data and returns `undefined` for a valid response | That is `status: 'success'` with `data === undefined` — a case outside the status × fetchStatus grid that the one-bit contract also loses; assert on `status`, not on the value | +| The paused state is unreachable because `networkMode: 'always'` is set | Drop the paused row for that query and state the mode in the component's contract, so a later mode change re-opens the row deliberately | +| The cache in use is SWR, RTK Query, or Apollo rather than TanStack Query | Map its fields onto the two axes before applying the table — where a cache exposes no separate fetch axis, build the explicit union of step 3 from the fields it does expose, and keep the pre-request case distinct from loading | +| Several queries feed one view | Combine on the axes, not the values: pending if any is pending, paused if any is paused — a merged `data === undefined` cannot distinguish them | + +## Instead of + +| If you are about to | Do this instead | Why | +| ------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Type the child's prop as `data \| undefined` and read `undefined` as "loading" | Pass `status` and `fetchStatus`, or an explicit state union built at the query boundary | A disabled or paused query is `pending` with `data === undefined`, `isLoading === false` and `isError === false`, so the child renders a spinner that no fetch will ever resolve | +| Show the spinner on `isPending` | Show it on `isLoading` (`isPending && isFetching`) and render the idle case separately | A lazy query is `pending` from the first render, so `isPending` puts a spinner on a query that was never requested | +| Add a timeout that turns a long spinner into an error | Render the `pending`/`idle` and `pending`/`paused` cells | The spinner is not slow, it is terminal — a timeout converts a missing state into a wrong one | +| Test the component by mocking the hook with `{data: undefined, isLoading: true}` and `{data: X}` | Drive one case per row of the step-2 table | Hand-written mocks reproduce the author's model of the states, so the combination that causes the bug is the one never constructed | + +## Sources + +- https://tanstack.com/query/latest/docs/framework/react/guides/queries — the `status` values (`pending` "The query has no data yet", `error`, `success`) and `fetchStatus` values (`fetching`, `paused` "The query wanted to fetch, but it is paused", `idle`); "Background refetches and stale-while-revalidate logic make all combinations for `status` and `fetchStatus` possible"; "The `status` gives information about the `data` … The `fetchStatus` gives information about the `queryFn`" +- https://tanstack.com/query/latest/docs/framework/react/reference/useQuery — `isLoading` "Is `true` whenever the first fetch for a query is in-flight. Is the same as `isFetching && isPending`"; `data` "Defaults to `undefined`"; `status` is `pending` "if there's no cached data and no query attempt was finished yet" +- https://tanstack.com/query/latest/docs/framework/react/guides/disabling-queries — a disabled query with no cached data is "status === 'pending' and fetchStatus === 'idle'"; "Lazy queries will be in `status: 'pending'` right from the start because `pending` means that there is no data yet … you likely cannot use this flag to show a loading spinner"; the `skipToken`/`refetch` incompatibility +- https://tanstack.com/query/latest/docs/framework/react/guides/network-mode — "Queries can be in `state: 'pending'`, but `fetchStatus: 'paused'` if they are mounting for the first time, and you have no network connection"; "it might not be enough to check for `pending` state to show a loading spinner" +- Source verification 2026-08-07 (`@tanstack/query-core@5.100.14`, `build/modern/queryObserver.js`): line 308 `const isPending = status === "pending"`, line 310 `const isLoading = isPending && isFetching`, line 332 `isPaused: newState.fetchStatus === "paused"` — the shipped derivation matches the reference, so `pending` + non-`fetching` yields `isLoading === false` with `data === undefined` diff --git a/wiki/frontend/index.md b/wiki/frontend/index.md index e6556c6..3b2642b 100644 --- a/wiki/frontend/index.md +++ b/wiki/frontend/index.md @@ -35,6 +35,7 @@ Match your situation to a "load when" line; load only matching pages. |------|-----------| | [race-conditions](data-fetching/race-conditions.md) | Repeated fetches with changing params can overlap (search-as-you-type, rapid tab/filter switches); UI intermittently shows results for a previous input; mutations race refetches | | [async-ui-states](data-fetching/async-ui-states.md) | Building any view backed by async data; users see blank screens, eternal spinners, or dead-end errors; reviewing loading/error/empty handling in UI code; deciding on skeletons vs spinners, retry affordances, empty states, background-refresh indication, or optimistic updates | +| [query-state-vs-fetch-state](data-fetching/query-state-vs-fetch-state.md) | Defining what a component receives from a server-state cache (TanStack Query and equivalents) and about to treat `data === undefined` as "loading"; a view shows a permanent spinner with no error and no retry; the query can be disabled (`enabled: false`, `skipToken`) or paused by the network mode; deciding what the presentational component's state prop should be | | [infinite-scroll](data-fetching/infinite-scroll.md) | Implementing infinite scroll or a load-more feed; an existing feed loses scroll position on back-navigation, duplicates/skips items, or spams page requests; choosing between infinite scroll and a load-more button | ## performance diff --git a/wiki/infrastructure/agent-orchestration/control-signals-vs-primary-artifacts.md b/wiki/infrastructure/agent-orchestration/control-signals-vs-primary-artifacts.md index 70633d6..cfe0e86 100644 --- a/wiki/infrastructure/agent-orchestration/control-signals-vs-primary-artifacts.md +++ b/wiki/infrastructure/agent-orchestration/control-signals-vs-primary-artifacts.md @@ -8,8 +8,8 @@ sources: - https://man.openbsd.org/tmux - https://code.claude.com/docs/en/hooks - https://pubs.opengroup.org/onlinepubs/9699919799/utilities/V3_chap02.html -last_verified: 2026-08-05 -related: [infrastructure-agent-orchestration-shared-run-state, infrastructure-agent-orchestration-pane-delivery-confirmation, infrastructure-agent-orchestration-session-completion-gates, platforms-shells-command-text-inspected-before-execution, platforms-tools-harness-mediated-tool-results, debugging-methodology-hypothesis-testing] +last_verified: 2026-08-06 +related: [infrastructure-agent-orchestration-shared-run-state, infrastructure-agent-orchestration-pane-delivery-confirmation, infrastructure-agent-orchestration-session-completion-gates, platforms-shells-command-text-inspected-before-execution, platforms-tools-harness-mediated-tool-results, debugging-methodology-hypothesis-testing, infrastructure-agent-orchestration-unattended-worker-questions, debugging-methodology-probe-path-vs-operation-path] --- # Deciding a Worker Is Done, Alive, or Dead from a Status File or Watcher Verdict @@ -67,6 +67,8 @@ progress by running a script that writes outside its worktree. | An editor/Write tool succeeds where Bash was refused for the same path | The guardrail inspects Bash command text only, so the two channels disagree by design. Do not use the working channel to route around the rule — report it | | Heartbeat is fresh but no commits in N intervals | Fresh heartbeat plus unchanged primary artifact is the stalled state, distinct from alive-and-progressing and from dead; handle it as its own case | | The monitor is the only thing that can see the worker | Add the primary-artifact check to the monitor rather than trusting its verdict; a monitor with no artifact check cannot produce evidence | +| Several workers go quiet almost simultaneously: liveness checks pass, heartbeats stay fresh, diffs stop growing | Before restarting anything, search each worker's pane/terminal tail for the CLI's usage-limit marker (e.g. `You've hit your session limit · resets HH:MM`) — a usage-limit pause is idle waiting, not a crash, so every liveness probe passes. After the stated reset time, send a resume prompt that orders: re-verify state (`git status`, rerun the tests) → remaining definition-of-done → completion signal. A bare "continue" sent before the reset is consumed by the same limit message, and a resume without the state re-check makes the worker guess where it stopped | +| The next task is dispatched to the same terminal immediately after the worker's done signal and fails runtime-unavailable | The done message is the worker's report time, not the substrate's release time — the CLI is still tearing down its stop-hook chain and the previous dispatch still occupies the terminal. Wait for the substrate's own idle signal (e.g. `orca terminal wait --for tui-idle`) before dispatching; and when the failed dispatch consumed the task, create a new task from the same spec — the consumed one cannot be retried | ## Instead of @@ -82,3 +84,4 @@ progress by running a script that writes outside its worktree. - https://man.openbsd.org/tmux — `has-session`, `display-message` behavior and session naming - https://code.claude.com/docs/en/hooks — hook architecture and path guardrails - https://pubs.opengroup.org/onlinepubs/9699919799/utilities/V3_chap02.html — exit codes and output redirection +- Field evidence 2026-08-06 (Claude CLI workers under tmux/Orca orchestration): three workers paused simultaneously on one usage-limit reset — identical `You've hit your session limit · resets 01:10` marker in each terminal tail while every liveness check passed; a state-recheck resume prompt sent after the reset resumed all three exactly at their interrupted step (test re-run). Same run: two dispatches issued immediately after `worker_done` both failed `runtime_unavailable` and consumed their tasks; dispatches issued after a `tui-idle` wait succeeded first try diff --git a/wiki/infrastructure/agent-orchestration/dispatching-after-a-completion-report.md b/wiki/infrastructure/agent-orchestration/dispatching-after-a-completion-report.md new file mode 100644 index 0000000..af9e5d6 --- /dev/null +++ b/wiki/infrastructure/agent-orchestration/dispatching-after-a-completion-report.md @@ -0,0 +1,84 @@ +--- +id: infrastructure-agent-orchestration-dispatching-after-a-completion-report +domain: infrastructure +category: agent-orchestration +applies_to: [orca, general] +confidence: verified +sources: + - "Orca CLI bundled skill guide: `orca skills get --topic orchestration --full` (app 1.4.177, command schema v1)" + - "`orca orchestration worker-start --help`, `orca terminal wait --help` (app 1.4.177)" +last_verified: 2026-08-12 +related: [infrastructure-agent-orchestration-control-signals-vs-primary-artifacts, infrastructure-agent-orchestration-pane-delivery-confirmation, infrastructure-agent-orchestration-session-completion-gates, infrastructure-agent-orchestration-shared-run-state, platforms-processes-non-interactive-cli-invocation] +--- + +# Reusing a Worker's Terminal for the Next Task After It Reports Completion + +## When this applies + +A worker reported completion (`worker_done`, a status write, an exit message) and +the orchestrator wants to hand that same terminal or runtime slot its next task. +Also when a start/dispatch call fails with a runtime-unavailable-class error +moments after a completion report, or when a task reaches a terminal `failed` +status without any worker having worked on it. + +## Do this + +1. **Treat the completion report as a claim about the work, not about the + runtime.** The report settles the task and the dispatch; the terminal is + released by a separate call. Between the two, the terminal is still owned by + the finishing dispatch and rejects a new one. +2. **Gate the next start on an idle check of the target terminal**, then start: + + ```sh + orca terminal wait --terminal "$HANDLE" --for tui-idle --timeout-ms 60000 --json + orca orchestration worker-start --task "$NEXT_TASK" --terminal "$HANDLE" --json + ``` + + The guide states this directly: "Wait for `tui-idle` before dispatching." +3. **Decide each settled dispatch's next owner before waiting again**, so no + terminal is left in the ambiguous middle state: + +| After an accepted completion report | Do | +|-------------------------------------|----| +| The same agent has an immediate follow-up task | Read `worker.agent_terminal_handle` from `worker-show --dispatch --json`, then `worker-start --task --terminal ` — this transfers cleanup ownership to the new dispatch | +| No follow-up for that agent | `worker-release --dispatch ` | +| The user asked to keep the terminal live for debugging | `worker-retain --dispatch `, and release it later | + +4. **Read the failed start's receipt instead of retrying it.** `worker-start` + exits 0 only for `ready`; a failed or `outcome_unknown` start exits nonzero and + returns `stage`/`failedStage`, `effects`, `residualResources`, and recovery + commands. Fix the stage the receipt names, then start again. +5. **Retry the same task through a linked replacement dispatch**, naming + placement explicitly, because retry does not inherit it: + `worker-start --task --retry-of --terminal ` + (or `--on`/`--worktree` plus `--agent`). +6. **Count the failures against the task.** After 3 consecutive failures on one + task the dispatch context circuit-breaks and the task is marked `failed`. Blind + retries spend that budget on the same unmet precondition; an idle check spends + none of it. + +## Edge cases + +| Case | Then | +|------|------| +| The completion report arrives but the terminal never reaches idle | This is the settled-dispatch/live-terminal state, not a stall — hold the handle and let the release path own it; do not close the terminal to force it | +| The start receipt says `outcome_unknown` | `worker-stop --dispatch ` and inspect again, or `worker-abandon --dispatch ` while accepting that resources may still be live — abandon performs no remote, process, or filesystem action | +| The task already reached `failed` from the circuit breaker | Recovering it means an explicit `task-update`, or a new task carrying the same spec; a `--retry-of` dispatch does not un-fail a circuit-broken task | +| The target is a bare shell rather than an agent CLI | Omit `--inject`, dispatch for tracking, and send the prompt with `terminal send --text … --enter`; the idle gate still applies | +| `worker-release` returns `release_pending` or `release_unknown` | Follow the recovery action in the receipt; substituting `terminal close` closes a terminal whose ownership the orchestrator has not proven | +| The idle wait times out on a long-running agent | A timeout is a checkpoint, not a failure — coding tasks run 15–60 minutes; keep waiting rather than starting a competing dispatch | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Start the next task the moment the completion report lands | Run the terminal idle check first, then start | The report settles the task; the previous dispatch still owns the terminal, so the start fails on an occupied runtime | +| Re-run the identical `worker-start` after it failed | Read the receipt's `stage`/`effects`/`residualResources`, fix that, then start with `--retry-of ` | Three consecutive failures on one task circuit-break it into `failed`, so a retry loop destroys the task it was meant to rescue | +| Let a `--retry-of` replacement pick its own placement | Repeat the intended `--on`/`--worktree` and `--agent`/`--terminal` choice on the retry | Retry links the attempt for provenance and deliberately does not inherit placement | +| Close the terminal yourself to free it for the next task | Transfer it with `worker-start --terminal ` or hand it to `worker-release` | The release path preserves inspectable output first and closes only the exact terminal the settled dispatch owns | + +## Sources + +- Orca CLI bundled skill guide, `orca skills get --topic orchestration --full` (app 1.4.177, command schema v1) — "Wait for `tui-idle` before dispatching"; "After processing each accepted `worker_done`, choose the terminal's next owner before you acknowledge the Delivery or wait again… run `orca orchestration worker-start --task --terminal --json` so Orca transfers cleanup ownership to the new Dispatch. Otherwise run `orca orchestration worker-release --dispatch --json`"; "After 3 consecutive failures on one task, the dispatch context circuit-breaks and the task is marked failed"; "It proves `failed` or `stopped`: start a replacement with `worker-start --task --retry-of ` plus an explicit `--on`/`--worktree` and `--agent`/`--terminal` choice. Retry does not silently inherit placement"; "Treat a `check --wait` timeout or `{count:0}` as a checkpoint, not a worker failure" +- `orca orchestration worker-start --help` (app 1.4.177) — "The call exits 0 only for ready. Failed or outcome_unknown exits 1 and JSON includes stage/failedStage, setup, effects, residualResources, and recovery commands when needed"; `orca terminal wait --help` — `--for exit|tui-idle`; `orca orchestration task-update --help` — statuses `pending, ready, dispatched, completed, failed, blocked` +- Field observation 2026-08-06 (dev-loop orchestration run, two occurrences): follow-up dispatches issued to a worker's terminal immediately after its completion report both failed on an unavailable runtime and consumed a dispatch attempt; the same start succeeded on the first try once the terminal was confirmed idle first diff --git a/wiki/infrastructure/agent-orchestration/pane-delivery-confirmation.md b/wiki/infrastructure/agent-orchestration/pane-delivery-confirmation.md index 077bf05..b9ce132 100644 --- a/wiki/infrastructure/agent-orchestration/pane-delivery-confirmation.md +++ b/wiki/infrastructure/agent-orchestration/pane-delivery-confirmation.md @@ -7,8 +7,8 @@ confidence: verified sources: - https://man7.org/linux/man-pages/man3/termios.3.html - https://man7.org/linux/man-pages/man1/tmux.1.html -last_verified: 2026-08-05 -related: [platforms-shells-option-like-argument-values, infrastructure-agent-orchestration-session-completion-gates, platforms-processes-non-interactive-cli-invocation] +last_verified: 2026-08-12 +related: [platforms-shells-option-like-argument-values, infrastructure-agent-orchestration-session-completion-gates, platforms-processes-non-interactive-cli-invocation, infrastructure-agent-orchestration-unattended-worker-questions] --- # Confirming a Keystroke Sent to a Terminal Pane Was Actually Consumed @@ -60,6 +60,12 @@ or escalate. | The send is a multi-line prompt | Send the body and the submit key as separate calls and check the indicator between them; a single blob can be consumed partially | | Several sends are in flight to one pane | Serialize them — one outstanding send per pane, confirmed before the next; interleaved input is reordered by the tty buffer, not by your script | | No busy indicator exists in the target | Require the artifact check from the table; without either, the harness cannot distinguish queued from consumed | +| You are binding a *new* unit of work to a pane the previous unit just reported finishing | Check for the idle prompt before binding, not after. A worker's "done" message is a report emitted from inside its turn, not the end of it — completion hooks keep the pane busy for minutes afterwards, and a dispatch bound to a busy pane is consumed as a failed unit rather than queued | +| The bind fails with a stage naming the runtime as unavailable | The pane is occupied by the tail of a turn — wait for the idle prompt and bind a fresh unit; the failed one is spent and is not retried in place | +| The bind fails with a stage naming the agent as unconfigured | The agent process behind the pane is dead even though the pane still renders and its worktree still resolves — no wait recovers it. Close the pane and create a new agent in worker mode, then bind | +| A bind is rejected for a pane/worktree mismatch | Pass the worktree alongside the pane on every bind; a pane identifier alone resolves against the coordinator's own checkout | +| The pane shows the prompt collapsed into a paste placeholder (`❯ [Pasted text #3]`) with no busy marker | The body arrived as one bracketed-paste block and the submit key was consumed with it — send `Enter` as its own `send-keys` call and re-read; [platforms-processes-non-interactive-cli-invocation] owns the paste mechanism | +| A send helper reports a queued outcome and its own follow-up wait then reports pick-up | That pair is a confirmation: the wait observed the target take the input. A helper that reports delivery without a wait has observed only the write | ## Instead of @@ -67,10 +73,14 @@ or escalate. |---------------------|-----------------|-----| | Diff `capture-pane` before/after and call a difference "delivered" | Check the busy/queued indicator first and use the diff only when it is absent | The tty echoes typed characters while the program is busy, so the diff reports success for the queued case the check exists to catch | | Sleep a fixed interval after `send-keys` and continue | Poll the indicator (or the artifact) until it clears, with a deadline | The right interval is the target's work time, which is what you are trying to measure | +| Read a send wrapper's success word or exit 0 as "the prompt is running" | Read it as "the keys reached the pane", then confirm submission from the pane: an empty input line plus the target's working indicator | The wrapper checks its own write, which succeeds whether the target submitted the input or parked it as an unsubmitted paste; the gap surfaces only as a phase timeout much later | | Resend on the first unchanged capture | Distinguish "busy" from "not delivered" before resending | Resending into a busy pane queues a duplicate that runs when the pane drains | +| Treat every failed bind the same way and retry it | Branch on the stage the failure names: wait-and-rebind for an occupied runtime, replace the agent for a dead one | The two look identical from outside — the pane renders in both cases — and retrying a dead agent spends units without ever succeeding | ## Sources - https://man7.org/linux/man-pages/man3/termios.3.html — `ECHO` in `c_lflag`: "Echo input characters." The terminal driver echoes independently of when the program calls `read()` - https://man7.org/linux/man-pages/man1/tmux.1.html — `send-keys` writes keys into a pane's input; `capture-pane` copies the pane's visible contents — neither reports whether the foreground process consumed the input +- Field observations 2026-08-06 (dev-loop 1.4.0 orchestrate, three dispatches): binding to a pane whose worker had just reported done produced a unit that went pending then failed and had to be recreated; a second pane rendered normally but its agent was dead, reported as an unconfigured-agent stage, and recovered only by closing the pane and creating a new worker-mode agent; a third bind was rejected for a pane/worktree mismatch and succeeded once the worktree was passed with the pane — the shipped script carries the same rule in `skills/orchestrate/scripts/orca-worker-start.sh` ("rejects the pair with `terminal_worktree_mismatch` (verified live)"), and `skills/orchestrate/SKILL.md` documents that a failed unit is replaced rather than retried in place +- Field observation 2026-08-12 (dev-loop orchestrate, 3 tmux worker sessions): `send-prompt.sh` returned 0/"delivered" for two workers whose panes both sat at `❯ [Pasted text #3]`/`#4` with the prompt unsubmitted, while the third returned "queued" and its follow-up `wait` reported pick-up — that one had actually submitted. Sending `Enter` as a separate key event to each stuck pane started both workers immediately - Field reproduction 2026-08-05 (tmux 3.7b, macOS): a pane running `sleep 6` received `echo SECOND_PROMPT_MARKER`. Pane content changed (diff = YES) and the marker appeared once as echoed text, while the command's own output line count stayed 0; after the sleep drained, the command ran and the output line appeared diff --git a/wiki/infrastructure/agent-orchestration/session-completion-gates.md b/wiki/infrastructure/agent-orchestration/session-completion-gates.md index 08e93fc..841d2a7 100644 --- a/wiki/infrastructure/agent-orchestration/session-completion-gates.md +++ b/wiki/infrastructure/agent-orchestration/session-completion-gates.md @@ -7,6 +7,9 @@ confidence: verified sources: - https://code.claude.com/docs/en/hooks last_verified: 2026-08-05 +related: [infrastructure-agent-orchestration-pane-delivery-confirmation, infrastructure-agent-orchestration-worktree-isolated-workers, platforms-processes-tool-diagnostics-without-a-failing-exit-code, infrastructure-agent-orchestration-dispatching-after-a-completion-report] + - https://csf.tools/reference/nist-sp-800-53/r5/ac/ac-5/ +last_verified: 2026-08-13 related: [infrastructure-agent-orchestration-pane-delivery-confirmation, infrastructure-agent-orchestration-worktree-isolated-workers, platforms-processes-tool-diagnostics-without-a-failing-exit-code] --- @@ -17,7 +20,8 @@ related: [infrastructure-agent-orchestration-pane-delivery-confirmation, infrast You are writing a completion gate — a `Stop`/`SubagentStop` hook or equivalent — that refuses to let an orchestrated worker session end while its recorded phase says the work is unfinished. Also when such a gate fires on a worker that did -exactly what its own prompt told it to do. +exactly what its own prompt told it to do — including when you *are* that worker, +parked at an instructed pause and receiving the nudge every turn. ## Do this @@ -48,6 +52,20 @@ exactly what its own prompt told it to do. 5. **Say what to do, not that something is wrong.** The block message names the phase, the next action, and the exact command that records completion — a blocked session's only input is that text. +6. **On the receiving side, hold the phase and report the mismatch.** When the + gate fires on you at a pause your own prompt instructed, keep the recorded + phase and tell the coordinator the gate's terminal set disagrees with the + prompt. Write only phases your role is authorized to write: + +| Phase | Written by | Because | +|-------|-----------|---------| +| `plan_ready`, `impl_done` | the worker | they record *its* progress and claim nothing about review | +| `approved`, `merged` | the coordinator only | they are the review verdict — the worker is the reviewed party | +| `done` | the worker, but only after the coordinator's approval message | it means "committed", which the approval authorizes | + + Silencing the gate by advancing the phase is the reviewed party issuing its + own approval; a downstream scheduler that treats those phases as dependency- + satisfying then dispatches work against an interface nobody reviewed. ## Edge cases @@ -58,6 +76,8 @@ exactly what its own prompt told it to do. | Several workers share one status directory | Match the entry by the session's resolved physical `cwd`; on macOS resolve `/var`→`/private/var` and symlinks on both sides before comparing | | A phase means "waiting on another worker" | Terminal — the worker cannot progress it; the orchestrator's wait loop owns that transition | | The worker cannot reach a terminal phase because the task is genuinely blocked | Provide a `failed` transition it may record itself; without one, the only escapes are fabricated completion or an eight-block override | +| You are the worker and the nudge repeats every turn at an instructed pause | Read it as a gate-vs-prompt mismatch, not as work you skipped; inventing extra work to satisfy it writes code the brief did not ask for | +| The gate's terminal set and the phase vocabulary live in different files | Cite both line numbers in the report — the fix belongs in the gate, and the coordinator is the one who can change it | ## Instead of @@ -66,8 +86,11 @@ exactly what its own prompt told it to do. | List only "success" phases as terminal | Add every phase at which the protocol instructs a stop, including mid-workflow handoffs | The gate otherwise fights the prompts the system issues, pushing the worker to fabricate completion or to do work it was told to hold | | Treat an unrecognized phase value as unfinished | Treat it as terminal and log the value | A typo or a newly added phase would otherwise trap sessions until someone reads the hook | | Rely on the block message alone to stop a loop | Return early on the harness's re-entry flag first | The message does not bound repetition; the flag is what makes the gate fire once | +| Advance your phase to a terminal value to stop a gate firing on you | Hold the instructed phase and report the gate-vs-prompt mismatch to the coordinator | The terminal values that would silence it are the review verdict; writing one makes the reviewed party its own approver, and the scheduler reads it as reviewed | ## Sources - https://code.claude.com/docs/en/hooks — `Stop`/`SubagentStop` input includes `stop_hook_active`; hooks check it and exit early to allow the stop. Claude Code overrides a Stop hook after it blocks eight times in a row without progress (cap adjustable via `CLAUDE_CODE_STOP_HOOK_BLOCK_CAP`) +- https://csf.tools/reference/nist-sp-800-53/r5/ac/ac-5/ — NIST SP 800-53 r5 AC-5: "Separation of duties addresses the potential for abuse of authorized privileges and helps to reduce the risk of malevolent activity without collusion. Separation of duties includes dividing mission or business functions and support functions among different individuals or roles" — the phase that records a review verdict belongs to the reviewing role, not the reviewed one +- Field reproduction 2026-08-13, dev-loop repo at `fa89dc2`: a worker parked at `impl_done` per `skills/orchestrate/templates/session-prompt.md:75` ("run `… status-update.sh {TASK} impl_done …` and wait") received "verification loop incomplete" on every turn, because `hooks/loop-gate.sh:55` accepts only `done|approved|merged|failed|""` while `skills/orchestrate/scripts/status-update.sh:6` lists `impl_done` as a first-class phase. The three values that would have silenced it are exactly the three `skills/orchestrate/scripts/ready-set.sh:74` counts as dependency-satisfying (`approved|merged|done`), and that file states the rule the fabrication would break: "A dependency counts as satisfied only at `approved` or higher, NOT at impl_done: a task that consumes an unreviewed interface has to be redone when rework changes that signature" - Field reproduction 2026-08-05, dev-loop repo at `95cf947`: `hooks/loop-gate.sh:55` lists `done|approved|merged|failed|""` as terminal, while `skills/orchestrate/templates/session-prompt.md:20` instructs a plan-phase worker to record `plan_ready` and "wait for an approval message. Do NOT write implementation code yet." A worker that followed its prompt exactly was blocked; the `stop_hook_active` early return at line 30 is what kept the block from repeating diff --git a/wiki/infrastructure/agent-orchestration/shared-run-state.md b/wiki/infrastructure/agent-orchestration/shared-run-state.md index e026b0d..66a69d1 100644 --- a/wiki/infrastructure/agent-orchestration/shared-run-state.md +++ b/wiki/infrastructure/agent-orchestration/shared-run-state.md @@ -7,8 +7,8 @@ confidence: field-tested sources: - https://git-scm.com/docs/git-worktree - https://man.openbsd.org/tmux -last_verified: 2026-08-05 -related: [infrastructure-agent-orchestration-control-signals-vs-primary-artifacts, infrastructure-agent-orchestration-session-completion-gates, infrastructure-agent-orchestration-worktree-isolated-workers, backend-common-concurrency-distributed-locks, backend-common-jobs-scheduled-job-overlap] +last_verified: 2026-08-17 +related: [infrastructure-agent-orchestration-unattended-worker-questions, infrastructure-agent-orchestration-control-signals-vs-primary-artifacts, infrastructure-agent-orchestration-session-completion-gates, infrastructure-agent-orchestration-worktree-isolated-workers, backend-common-concurrency-distributed-locks, backend-common-jobs-scheduled-job-overlap, testing-data-test-data-and-isolation, infrastructure-agent-orchestration-pane-delivery-confirmation] --- # Orchestration State Kept in a Shared Directory Inside the Repository @@ -67,6 +67,7 @@ running. | A foreign run merged your task branches without your gate | Stop and reconstruct from `git reflog` and the default branch's history before continuing; the integration branch no longer reflects only your approvals | | Worktrees were removed but their branches remain | `git worktree prune` clears the administrative files; the branches persist and still signal a prior run — read branch names for the run id | | Run ids are generated from a timestamp at one-second resolution | Two runs starting in the same second collide; add a random suffix or the process id | +| The coordinator re-delivers a prompt to a stalled worker and also re-seeds that task's status file | Reset the status **before** sending the prompt, never after — the file is last-write-wins, and the worker's first progress signal can land between your send and your reset. The overwrite leaves the watcher waiting on a phase the worker already left while the worker waits for the next instruction: a deadlock with both sides idle | ## Instead of @@ -76,8 +77,11 @@ running. | Treat an unfamiliar task id in the state directory as leftover junk | Check worktrees, recent branches, and the default branch HEAD first | Stale state and a live concurrent run are indistinguishable from the files alone | | Have your watcher act on every status file it sees | Filter to the task ids this run created | Otherwise another run's completion signal reads as your own task finishing | | Continue after noticing the default branch moved | Identify what merged, then decide | Your integration branch may already be missing or duplicating merged work | +| Re-seed a task's status to `pending` after delivering a re-prompt | Reset first, then send, and let the worker's next write own the file | Two uncoordinated writers on one status file resolve by last write; the coordinator's late reset erases the worker's progress signal and the pane log and the file then disagree | ## Sources - https://git-scm.com/docs/git-worktree — linked worktrees share one repository with separate working directories - https://man.openbsd.org/tmux — session and pane management for substrate-level coordination +- https://en.wikipedia.org/wiki/Race_condition — programs colliding on a shared file produce order-dependent results; coordination (locking or a single writer) is required for a deterministic outcome +- Field evidence 2026-08-17 (linkly run iss0817, task t60): the worker's `plan_ready` write (12:53:0x) was overwritten by the coordinator's `pending` re-seed (12:53:07) issued after the re-prompt was sent — the tmux pane recorded "status set to plan_ready" while the file read `pending`, and the watch waited on a phase the worker had already passed. Moving the reset before the send removes the window diff --git a/wiki/infrastructure/agent-orchestration/unattended-worker-questions.md b/wiki/infrastructure/agent-orchestration/unattended-worker-questions.md new file mode 100644 index 0000000..0a0fbbd --- /dev/null +++ b/wiki/infrastructure/agent-orchestration/unattended-worker-questions.md @@ -0,0 +1,89 @@ +--- +id: infrastructure-agent-orchestration-unattended-worker-questions +domain: infrastructure +category: agent-orchestration +applies_to: [tmux, general] +confidence: verified +sources: + - https://github.com/anthropics/claude-code/issues/50728 + - https://github.com/anthropics/claude-code/issues/29530 + - https://man7.org/linux/man-pages/man1/tmux.1.html +last_verified: 2026-08-08 +related: [infrastructure-agent-orchestration-usage-limit-paused-workers, infrastructure-agent-orchestration-pane-delivery-confirmation, infrastructure-agent-orchestration-control-signals-vs-primary-artifacts, infrastructure-agent-orchestration-shared-run-state, infrastructure-agent-orchestration-session-completion-gates, platforms-processes-non-interactive-cli-invocation] +--- + +# A Worker Agent Asks a Question With No Human at Its Terminal + +## When this applies + +An orchestrator runs agent workers unattended (tmux panes, a task runner, an SDK +subprocess) and one worker raises a question through its own in-band question UI — +a numbered chooser, a confirmation screen, a trust or re-auth prompt. Also when a +worker is judged stalled with a live terminal and no task-level error, or when a +worker reports a decision it "assumed" that nobody was asked about. + +## Do this + +1. **Treat an in-band question as unanswerable and design an out-of-band channel + for it.** Give the worker one command that writes a durable question record the + coordinator polls — `{ts, taskId, question, options, worktree}` in a + `questions/` directory beside the run's status directory + ([infrastructure-agent-orchestration-shared-run-state]) — and have the worker + end its turn after writing it. The record survives the worker's turn; a UI + waiting for a keystroke does not. +2. **State the channel in the worker's first prompt**, naming the command and the + rule that a decision needing the coordinator is written to that channel rather + than raised locally. A worker follows the prompt it was given; without the rule + it reaches for its default question UI. +3. **Make a pending question a distinct wake reason** in the watcher, separate + from "failed" and from "all tasks reached the phase", so it is handled and + cleared rather than aggregated into a timeout. +4. **Classify before acting on a stall — read the terminal tail first.** "Wedged + on a question" and "finished but never reported" are identical from outside and + need opposite responses: + +| Terminal tail shows | Do | +|---------------------|----| +| A question UI with selectable options | Unblock it by key (step 5), then re-drive the interrupted work (step 6) | +| The idle prompt, work visibly complete | Ask for the completion signal; the worker finished and skipped its report | +| The busy/working indicator | Not a stall — keep waiting ([infrastructure-agent-orchestration-pane-delivery-confirmation]) | +| A usage-limit or re-auth notice | Idle waiting, not a crash; resume after the stated reset, following [infrastructure-agent-orchestration-usage-limit-paused-workers] | + +5. **Unblock a question UI with an allowlisted key sequence, validated whole + before any key is sent.** Restrict the allowlist to navigation and answer keys + (`Up Down Left Right Enter Escape Tab Space 0-9 y n`), send one key per call so + ordering is deterministic and a failure names the key that did not land, and + re-read the terminal until the idle prompt returns — a chooser can have a + selection step and a confirmation step, and one key answers only the first. +6. **Re-send the prompt that was in flight when the question opened.** Text typed + or pasted into the input line before the UI opened is not submitted by the keys + that answer the UI; the turn ends quietly with the work never started. Confirm + from the artifact the prompt was supposed to produce, not from the terminal. + +## Edge cases + +| Case | Then | +|------|------| +| The worker runs with no TTY (SDK subprocess, container, CI) | The question tool does not block — it resolves immediately with empty answers and the agent continues as if answered, so the failure is a silent wrong decision instead of a stall. The out-of-band channel is the fix in both substrates | +| Liveness and heartbeat checks all pass while nothing progresses | A live PTY on a question UI is alive-and-not-progressing, a third state distinct from alive and dead; add a per-agent activity check to see it | +| The stall detector keys off terminal output timestamps | A TUI repaints its spinner continuously, so a wedged worker reports output "0 seconds ago" — key off the agent's own state timestamp instead | +| The out-of-band ask has a timeout and it expires | A timeout leaves the question pending; it is not an answer. Resume the same question rather than deciding it in the worker or asking it again | +| Answering the question requires a decision the coordinator also cannot make | Record it against the task and hand the whole record to the human once, rather than blocking each worker separately | +| The question record's task id is attacker- or environment-derived | It becomes a filename — reject `/`, `.`, `..`, and empty before writing, so a record cannot escape the questions directory | + +## Instead of + +| If you are about to | Do this instead | Why | +|---------------------|-----------------|-----| +| Let a worker raise its question through its own interactive UI | Have it write a durable question record and end its turn | The UI needs a human at that terminal; with a TTY it waits indefinitely, and without one it self-answers empty | +| Restart or replace a worker that a stall check flagged | Read the terminal tail and classify first | A worker waiting on a question is intact and one keystroke from resuming; restarting discards its completed work | +| Send free text to answer a chooser | Send allowlisted key events, validated as a set before the first is sent | A chooser reads key events, and a half-delivered sequence leaves it in a state neither side can name | +| Tell a worker to "decide it yourself and note the assumption" | Give it the question channel and have it wait for the answer | Measured on a 3-worker run: at 600 s and 900 s both workers instead proceeded on a conservative assumption and reported the guess after the fact | + +## Sources + +- https://github.com/anthropics/claude-code/issues/50728 — `AskUserQuestion` in a headless/no-TTY environment (Docker, `claude-agent-sdk` 0.1.63, bundled CLI 2.1.114): "auto-resolves immediately with empty answers", completing in ~37 ms before a `can_use_tool` callback or `PreToolUse` hook can intervene; the agent receives "User has answered your questions: ." and continues. Closed as not planned +- https://github.com/anthropics/claude-code/issues/29530 — same tool "does not render any interactive UI (question text, selectable options)" and returns an empty answer (CLI 2.1.63, open) — a second report that the in-band channel cannot be relied on to reach a human +- https://man7.org/linux/man-pages/man1/tmux.1.html — `send-keys` writes key events into a pane; named keys are sent without `-l`, which sends the literal characters instead +- Shipped implementation, dev-loop 1.4.2 `skills/orchestrate/`: `scripts/ask-coordinator.sh` writes one atomic `questions/.json` record per task and refuses a task name containing `/`, `.`, or `..`; `scripts/watch-status.sh` surfaces it as its own exit code within one poll; `scripts/send-prompt.sh keys` validates every key against the allowlist before sending any and sends one `send-keys` per key; `scripts/orca-worker-stalled.sh` documents the measurement behind the third state — three workers held a live PTY on an interactive prompt for 75 minutes with byte-identical diffs while every alive/dead check passed, and terminal-output timestamps were measured and rejected because a TUI's repaint keeps them fresh +- Field observation 2026-08-08 (dev-loop orchestrate, tmux worker `lo-4-qag1`): the worker raised a numbered chooser in its pane and sat 30 minutes past its silence threshold with the terminal alive. Two rounds of number-then-Enter cleared it (selection, then confirmation); the prompt queued in the input line before the chooser opened was never submitted, and the phase reached its next state only after that prompt was re-sent diff --git a/wiki/infrastructure/agent-orchestration/usage-limit-paused-workers.md b/wiki/infrastructure/agent-orchestration/usage-limit-paused-workers.md new file mode 100644 index 0000000..115f9ec --- /dev/null +++ b/wiki/infrastructure/agent-orchestration/usage-limit-paused-workers.md @@ -0,0 +1,87 @@ +--- +id: infrastructure-agent-orchestration-usage-limit-paused-workers +domain: infrastructure +category: agent-orchestration +applies_to: [claude-code, tmux, orca, general] +confidence: verified +sources: + - https://code.claude.com/docs/en/errors + - https://code.claude.com/docs/en/costs + - https://github.com/anthropics/claude-code/issues/5977 + - https://github.com/anthropics/claude-code/issues/36320 +last_verified: 2026-08-12 +related: [infrastructure-agent-orchestration-unattended-worker-questions, infrastructure-agent-orchestration-control-signals-vs-primary-artifacts, infrastructure-agent-orchestration-pane-delivery-confirmation, infrastructure-agent-orchestration-dispatching-after-a-completion-report, infrastructure-agent-orchestration-shared-run-state] +--- + +# Worker Sessions Paused by a Provider Usage Limit + +## When this applies + +Several agent workers billed to one account go quiet within minutes of each other, +their diffs stop, and every liveness check still passes. Also when one worker's +terminal tail carries a `You've hit your … limit · resets …` notice, or when you +are deciding whether to restart, replace, or wait on a worker that reports no +task-level error. + +## Do this + +1. **Read the terminal tail for the limit marker before classifying the stall**, + because which limit was hit decides whether waiting is the only move: + +| Tail shows | Do | +|------------|----| +| `You've hit your session limit · resets