Skip to content

feat: apply evidence-based confidence updates - #11

Merged
fabbrik merged 7 commits into
mainfrom
feat/3-4-confidence-updates
Sep 22, 2026
Merged

fabbrik merged 7 commits into
mainfrom
feat/3-4-confidence-updates

Conversation

@fabbrik

@fabbrik fabbrik commented Sep 22, 2026

Copy link
Copy Markdown
Owner

Summary

Story 3.4. Trust stops being a number stamped once at finalization.

  • The heuristic is (1 + S) / (2 + S + F), versioned, independent of completion score and of status — a number never makes an ineligible record eligible. Documented as a heuristic, not a calibrated probability.
  • Core owns the arithmetic. It reads the record, computes the counters and score, and submits with the revision it read. The adapter never derives a score.
  • Independence is keyed on (experience, run, round) for machine evidence and (experience, reviewer, run) for human evidence, enforced by a generated column and a partial unique index.
  • A duplicate key records its submission and changes nothing else — not the counters, status, revision or updated_at. That last one matters: recency ranking and expiry read UpdatedAt, so a moving timestamp would have kept a record artificially young forever.
  • A contradiction contests the record in the same transaction, and de-indexes it.
  • Counters and confidence move only with an event that recorded them, enforced by the projection trigger, so a direct rewrite that advances the revision is still refused.

The claim this story can and cannot make

The generated key stops a caller choosing the key string. It does not stop a caller choosing its inputs: RunId and VerificationRoundId are GUIDs the caller supplies, and this schema has no run or round table to bind them to. They are a host trust boundary, exactly like the reviewer identity — established from the host's own run bookkeeping and closed rounds, never passed through from agent output. The docs now say that instead of claiming inflation is impossible.

Review

Three reviewers raised 37 findings; 36 fixed, 1 deferred (the ledger's read, foreign key and retention, assigned to 4.5). The serious ones: the inflation claim above; an uncounted duplicate still refreshing recency and still contesting a record; and a guessed evidence ID returning another tenant's scores because the replay comparison had no scope predicate.

One item is documented rather than changed: the eligibility gate runs before the store's idempotency check, so retrying evidence after the record left eligibility reports Ineligible for an update that did commit. Nothing is written or lost; moving the gate needs the accepting-status set in Abstractions and a new outcome member. It is pinned by a test and stated on the outcome and in the README.

Test plan

  • dotnet build --configuration Release: 0 warnings, 0 errors
  • dotnet test --configuration Release: 879 of 879 pass (Core 466, Postgres 216, Vectors 54)
  • Container tests per matrix row: the 2/3 → 3/4 → 3/5 sequence, both duplicate kinds, reused and replayed evidence IDs, the concurrent race, a cross-scope probe, a revision-advancing rewrite refused with 42501, and de-indexing on a contest
  • CI green

🤖 Generated with Claude Code

fabbrik and others added 7 commits September 18, 2026 21:17
Add ExperienceFinalizationService: one call that loads a completed captured
run, evaluates its own closed verification round, checks host authorization and
the host's storage decision, reflects, creates the record as Candidate, and
commits its initial lifecycle event to Validated or Quarantined. Record,
reflection and event IDs derive from the run, so a retry re-derives them and
converges instead of duplicating.

Required checks can now name the evaluator kind that may satisfy them, and Core
and Storage.Postgres each expose a service-registration extension. The MAF
adapter finalizes a fully captured run through the host's resolver.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Add IExperienceCandidateSource with a PostgreSQL implementation: scope, status
and confidence filtering plus a full-text match over a generated tsvector
column added by migration 0003. Core's ExperienceRetrievalService applies
expiry and environment eligibility, then ranks candidates on relevance,
confidence, recency, status and environment compatibility with configurable
validated weights, exposing every normalized component and effective weight.

Retrieval is bounded by a timeout that returns an empty result with a timeout
signal rather than throwing, caller cancellation stays distinct, and a capped
candidate pool is reported through Truncated.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Add IExperienceEmbeddingIndex and a domain-typed embedding generator port,
with a new AgentExperience.Storage.Postgres.Vectors package implementing them
over plain Npgsql and Pgvector. Embeddings are derived data: indexing runs
after the canonical commit, embeds only the sanitized retrieval summary of an
eligible record, and writes conditionally on the record's exact revision, so a
stale write is rejected and a deleted record is never recreated.

Retrieval gains a vector channel merged with the text channel under the same
eligibility, timeout and ceiling. A model or dimension mismatch, an
unavailable provider, or a vector-channel failure produces an explicit
flagged text-only result rather than an incompatible comparison.

The vectors package owns and applies its own schema, so a text-only
deployment never runs CREATE EXTENSION vector.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Add ExperienceContextProvider, a simple-tier AIContextProvider the host adds
to ChatClientAgentOptions.AIContextProviders. It retrieves ranked experience
before an invocation, re-checks every candidate against the same eligibility
rules retrieval applies, asks the host's injection decision, and injects a
delimited Historical Reference carrying source, confidence, applicability and
an evidence summary -- never raw payload content, never a cut record.

Limits drop whole records and record every omission. Retrieval that is empty,
times out or fails leaves the agent running normally. Labeling marks the
content untrusted; the authorization boundary is what denies an unauthorized
tool call, and a test proves it does when the model obeys injected text.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Add IExperienceGrantStore and migration 0005: an administrator, with authority
the host supplies explicitly, grants one record to a recipient scope with a
reason and an expiry. The grant row and its audit event commit together, and a
unique partial index allows at most one active grant per recipient, so revoking
the grant an administrator knows about ends that recipient's access.

Reads widen in SQL only: get, text search and vector search match their exact
scope or an active grant, evaluated against the database clock. Writes,
lifecycle commits, history and enumeration stay owner-only. The read that
applied the predicate marks a record as shared, so Core and the injection
provider keep strict scope equality for everything else, the host's risk
policy can deny borrowed experience, and the model is told the lesson came
from another scope.

Deployments without the grants table, or with SELECT only on the record table,
degrade to exact-scope reads rather than failing.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Complete the MVP transition table (reinforce, contest, stale, supersede,
revoke), refuse same-state events, and add supersession with a recorded
replacement whose eligibility and cycle rules are re-decided inside the commit
transaction under row locks, so two concurrent supersessions cannot store the
cycle they would each individually pass.

Close the null-prior bypass: a first event may only record the status the
record is already in, enforced in Core and in the projection guard.

Migration 0006 makes the audit trail enforced rather than conventional:
statement- and row-level triggers, ENABLE ALWAYS so replication cannot skip
them, covering update, delete and truncate on both event logs, monotonic
revocation and expiry plus pinned identity on grants, and forward-only
revisions on the record projection. Constraints are NOT VALID with a
documented validate step, so an upgrade cannot abort on existing rows.

History is bounded and cursored and carries each event's stored timestamp and
applied revision. Leaving eligibility removes the record's embedding, which is
storage hygiene rather than a reachability boundary, and never fails the
transition.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Accepted evidence moves reuse confidence through the versioned heuristic
(1+S)/(2+S+F). Core reads the record, computes the counters and score, and
submits them with the revision it read; the adapter enforces independence with
a generated key and a partial unique index, and applies the ledger row, the
counters, any status change and the lifecycle event in one transaction.

Independence is keyed on (experience, run, round) for machine evidence and
(experience, reviewer, run) for human evidence. A duplicate key records its
submission and changes nothing else -- not the counters, the status, the
revision or the timestamp -- so replaying one observation can neither inflate
a score nor keep a record artificially recent.

The run and round are a host trust boundary, like the reviewer identity: the
generated key stops a caller choosing the key string, not its inputs, and the
docs now say so rather than claiming inflation is impossible.

A contradiction contests the record in the same transaction. Counters and
confidence move only with a lifecycle event that recorded them, enforced by
the projection trigger.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
@fabbrik
fabbrik merged commit d684016 into main Sep 22, 2026
1 check passed
@fabbrik
fabbrik deleted the feat/3-4-confidence-updates branch September 22, 2026 21:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant