fix: announce Keyman side and lock purge audit counts (v0.87.1) - #182
Closed
cursor[bot] wants to merge 120 commits into
Closed
fix: announce Keyman side and lock purge audit counts (v0.87.1)#182cursor[bot] wants to merge 120 commits into
cursor[bot] wants to merge 120 commits into
Conversation
Confirmed against real Milestone 2 SAP CRM VOC data: post_summary.py's
R&R extraction forced every named actor into a person slot, but real
business correspondence routinely names an organization acting in its
own name ("당사," "SEWA," "Siemens," "GECO"), not an individual.
- RoleResponsibility.actor_name (renamed from person_name) gains
actor_type_code (prov_person/prov_organization, W3C PROV-O grounded:
Lebo, Sahoo, & McGuinness, 2013) and an LLM-inferred
affiliated_organization_name for person actors -- a bare name
without an employer is hard to place.
- Ontology: :RoleActorPerson rdfs:subClassOf prov:Person,
:RoleActorOrganization rdfs:subClassOf prov:Organization -- genuine
subclasses of the real external PROV-O classes, distinct from the
ontology's existing :Person (a cataloged Keyman with a stable
person_id; an R&R actor is a free-text name with no cataloged
identity).
- migrations/0012_role_responsibility_agent_type.sql renames the
column via RENAME COLUMN (preserves existing rows), not a
drop/recreate.
- Popup R&R list shows a Person/Organization badge and the inferred
affiliation; only a person actor still links to the Keyman panel.
- Also fixes a real deployment gap found via browser E2E testing:
migrations 0005-0011 had accumulated on main without ever being
applied to the long-running demo Postgres volume, surfacing as
CORS-looking failures (missing-table 500s lose their CORS header)
on Evaluate, Reports, Summary, and Chat.
ADR 0006.
Co-Authored-By: Claude Sonnet 5 <[email protected]>
Drop real-organization names from docs, prompts, and comments. Seed a synthetic organization actor so the Person/Organization badge is visible without a live LLM, and reject unknown actor_type_code values.
PersonMention now carries an optional job_title extracted by the LLM from role phrasing (e.g. "our legal counsel, Sam Okonkwo"), not just named affiliations. cataloged_person.last_known_job_title persists it, and _upsert_person treats a conflicting stated title as evidence that a same-name match is a different real person rather than a re-mention, so two "Kim Cheolsu"s with different titles get distinct person rows. Keyman panel renders the title next to the person and per-affiliation role_title, which existed in the schema but was never surfaced before. Migration 0013 adds the column additively; 0001_initial_schema.sql bakes it in for fresh installs, matching this repo's existing pattern.
After make seed, Ada West / Priya Nair / Jordan Hale carry last_known_job_title so the new title chip is visible without a live extraction.
Strix flagged the local-dev password literal in seed_demo_data.py after this branch started editing that file. make seed still injects the compose default; a direct script run requires KEYCLOAK_ADMIN_PASSWORD.
…70.0)
Real post text named a company sub-unit ("설계팀"/design team) that
neither ADR 0006's prov_person nor prov_organization fits -- it's part
of a company, not a person and not the company itself. actor_type_code
gains prov_team, grounded in the W3C Organization Ontology's
org:OrganizationalUnit (Reynolds, 2014), a different W3C vocabulary
from PROV-O that exists specifically for this meso-level case.
A team actor requires affiliated_organization_name in the same way a
person actor does -- unlike an organization actor, a team's own name
never answers "which company." Fixed a real bug the new type surfaced:
the R&R badge's label text was a binary Person/Organization ternary
that would have mislabeled a team as "Organization" (the CSS class
name was already generic; the display text was not).
Migration 0014 is purely additive (one lookup row insert), no schema
change -- actor_type_code already stores an arbitrary FK'd code.
Real post text names organizations by abbreviation ("한수원" for
"한국수력원자력") that corporate_hierarchy_resolution's character-
similarity matching cannot bridge -- an initialism shares almost no
substring with its expansion, so no similarity threshold recovers it.
New lineageweave/organization_name_resolution.py: an LLM proposes the
full name from context (or declines with UNKNOWN), then the *existing*
relation_verification Searxng client cross-verifies the specific raw/
resolved pairing -- no second web-search integration built, reusing
what this repo already has for a structurally identical problem. Only
a search-corroborated resolution is ever substituted in for
resolve_corporate_entity; an unresolved or unverified name still flows
through unchanged, same never-trust-an-unverified-guess discipline as
every other channel here.
Cached in a new organization_name_resolution table
(migrations/0015), keyed by the raw name so the same abbreviation
across many posts is resolved once, not re-queried every mention.
Grounded in SKOS skos:altLabel/skos:prefLabel (Miles & Bechhofer, 2009).
Wired into backend/app/keyman_ingestion.py's affiliation loop and the
private real-data batch script's paced re-implementation of it -- which
was also found missing role_title persistence entirely (a stale copy
predating that feature), fixed alongside this.
Known, documented gap (ADR 0008): the same request's entity-
relationship classification step still uses the raw, unresolved
organization names -- not fixed here, tracked honestly instead of
silently shipped as if both sides already agreed.
…atch (v0.72.0) _parse_description required a single regex to match TEXT/CAPTION/TAGS in that exact order in one pass. Reproduced live against real embedded images from the Milestone 2 batch: real vision responses with the content right but the formatting only mostly right (bolded labels, reordered labels, a missing TAGS line) were rejected wholesale, producing the same "[image: content unavailable]" placeholder as a genuinely unconfigured vision channel -- discarding real, already-paid-for content, not a "genuinely could not get it" case. Each label is now parsed independently by scanning lines for a TEXT:/CAPTION:/TAGS: prefix (tolerant of markdown emphasis and any order); only a response with neither TEXT nor CAPTION content raises ImageDescriptionParseError. Multi-line TEXT (real multi-line OCR output) is still preserved with real newlines, not flattened.
…ity-agent-ontology # Conflicts: # CHANGELOG.md # frontend/package.json # lineageweave/__init__.py # pyproject.toml
…v0.74.0)
Extraction runs per-post; a team or organization's identity did not
survive across posts the way a Keyman's already did via
cataloged_person -- "설계팀" named in ten posts was ten unrelated
strings, not one entity the KG could link through. Extraction results
must themselves become cross-post lineage clues, not just per-post
artifacts.
New cataloged_team catalog (migrations/0016), identity key (team_name,
affiliated_organization_name) since a bare team name is not by itself
identifying ("설계팀" exists at many real companies) -- reuses the same
resolve_corporate_entity matching Keyman affiliations already use for
the team's parent org, not a second algorithm. An organization actor
resolves against the existing corporate_entity catalog directly, no
new table needed.
knowledge_graph_edges_for_post gains three new edge kinds
(edge_mention_team, edge_team_affiliation, edge_mention_organization)
as distinct object properties, not widened domain/range on the
existing :mentions (which would let RDFS entail every :mentions
subject is both a person and a team). persist_post_summary now
resolves each R&R actor's identity and calls the same
persist_edges_for_post Keyman ingestion already uses -- one function
computes a post's whole edge set regardless of trigger.
A person R&R actor is opportunistically joined to an existing
cataloged_person row by name, never originated by R&R itself --
documented as a real, deliberate gap in ADR 0009 (cataloged_person
needs person_side_code, which R&R's prompt does not currently ask
for), not silently half-done.
…y (v0.75.0) corporate_hierarchy_resolution's similarity matching only ever finds an ALREADY-cataloged corporate_entity -- it has no path to create one. Real Milestone 2 data confirmed the actual consequence: 0 of 4,154 person_affiliation rows and 0 of 9,852 R&R organization-actor mentions ever resolved, because corporate_entity for the real dataset only holds the employer's own 2-row hierarchy. The standing "통합 고객사 계열 tree AI" requirement (Samsung -> Samsung Electronics Korea -> ...) was never actually populated for real extraction. New lineageweave/corporate_hierarchy_inference.py: an LLM proposes a Group/Company/Plant placement (level + parent name) from the post's own text, or declines with UNKNOWN. New backend/app/corporate_entity_ingestion.py's get_or_create_corporate_entity tries similarity matching first (unchanged), then only creates a real new row once the proposal is corroborated by the *existing* relation_verification Searxng client -- no new search integration, reusing the same reused-verification-client pattern ADR 0008 already established. Recurses up a bounded (4-level) parent chain so the whole hierarchy gets real parent_entity_id links, not an orphaned row. Auto-created corporate_entity_code values are AUTO-<hash>-prefixed -- that column doubles as the real login corp-code Keycloak claim, so an auto-created counterparty must never collide with that namespace. Wired into both existing organization-resolution call sites (keyman_ingestion.py's affiliation loop, post_summary_ingestion.py's R&R organization-actor loop) rather than a third path, so both routes to corporate_entity share one creation policy. Found and fixed a pre-existing gap in backend/tests/test_api.py's seeded_db fixture along the way: it never seeded the 'plant' corporate_entity_level lookup row.
The fail-closed rollback script starts an explicit transaction. On an autocommit connection a RAISE left that transaction aborted, so the empty-registry cleanup could not run.
Buyer gap: after #95 the home Analysis runs row was inert text. Clicking the seeded Demo Corp lineage run now loads GET /api/analysis-runs/{id} and shows cutoff, requested date, and document count. Hidden runs stay not-visible. Synthetic aggregates only -- never a DSN or source SQL.
Buyer gap: after #100 the detail showed cutoff and counts but not the legal lifecycle the registry already stored. GET /api/analysis-runs/{id} now returns labeled status_history (Pending → Running → Succeeded with occurrence times). The list stays latest-status only. Hidden runs still 404 and never leak events. Failure codes stay machine tokens. Synthetic Demo Corp seed only.
Buyer gap: after #102 the run detail showed history but no way to open a post. Detail now lists ABAC-visible titles in the run's scope. Other-corp private posts stay hidden. List payloads stay aggregates-only. Synthetic titles only.
PR #91 landed an adaptive-orchestration ADR 0013 on the #74 base after this slice already used 0013 for the normalized analysis-run registry. Renumber the adaptive record to 0015 so ADR numbers stay unique. Co-authored-by: Seongho Bae <[email protected]>
The #74 changelog fold still called that decision ADR 0013. This stack keeps the analysis-run registry as ADR 0013, so the adaptive record is 0015. Co-authored-by: Seongho Bae <[email protected]>
Migration 0016 no longer deletes overlapping Keyman mention_context. Analysis-run detail lists only posts known at knowledge_cutoff. Keyman org enrichment finishes before the write transaction. Replace remaining real organization names with synthetic AGP examples. Co-authored-by: Seongho Bae <[email protected]>
* feat: seed a TEPP analysis run through tepp_client (v0.84.0) Buyer gap: home Analysis runs only showed lineage reconstruction. make seed now records a Demo Corp TEPP measurement via tepp_client. The default transport is unavailable, so the row is Failed / tepp_not_available -- never a fabricated theta. TEPP stays a wire client, not a local psychometric engine. * fix: fail-closed TEPP seed on the shared Demo Corp snapshot #111 still marked a live unused envelope Succeeded, named a different capture than the registry row, and re-inserted frozen counts. Seed now reuses the lineage snapshot (ADR 0013), skips count inserts after the first run, and keeps missing or unused TEPP Failed. The home list tells the operator to open the run and connect TEPP; detail history keeps tepp_not_available. Co-authored-by: Seongho Bae <[email protected]> --------- Co-authored-by: Seongho Bae <[email protected]> Co-authored-by: Cursor Agent <[email protected]> Co-authored-by: Seongho Bae <[email protected]>
The #89 review asked for 12-character code and config prefixes so an operator can match the approved revision. Full digests stay on the API only. Do not merge until this review item is checked. Co-authored-by: Cursor Agent <[email protected]> Co-authored-by: Seongho Bae <[email protected]>
* feat: seed a TEPP analysis run through tepp_client (v0.84.0) Buyer gap: home Analysis runs only showed lineage reconstruction. make seed now records a Demo Corp TEPP measurement via tepp_client. The default transport is unavailable, so the row is Failed / tepp_not_available -- never a fabricated theta. TEPP stays a wire client, not a local psychometric engine. * fix: fail-closed TEPP seed on the shared Demo Corp snapshot #111 still marked a live unused envelope Succeeded, named a different capture than the registry row, and re-inserted frozen counts. Seed now reuses the lineage snapshot (ADR 0013), skips count inserts after the first run, and keeps missing or unused TEPP Failed. The home list tells the operator to open the run and connect TEPP; detail history keeps tepp_not_available. Co-authored-by: Seongho Bae <[email protected]> * fix: keep failed-run next actions kind-specific A failed lineage row must not tell the operator to connect TEPP. Stacked PRs now run the same GitHub Checks as PRs to main. Co-authored-by: Seongho Bae <[email protected]> * docs: keep TEPP next-action copy off failed lineage rows Co-authored-by: Seongho Bae <[email protected]> * fix: keep TEPP corpus hint off a succeeded measurement A calibrated TEPP row must not tell the operator to replace Failed. Co-authored-by: Seongho Bae <[email protected]> --------- Co-authored-by: Seongho Bae <[email protected]> Co-authored-by: Cursor Agent <[email protected]> Co-authored-by: Seongho Bae <[email protected]>
….84.1) (#127) * fix(ui): keep analysis-run digests audible and warn on live posts aria-label on the digest paragraph hid the prefixes from assistive technology. Move the label to a group, keep prefixes as visible text, and put the full digest on hover. Tell the operator that a cutoff title opens the live body so they compare it with the run clock. Co-authored-by: Seongho Bae <[email protected]> * docs: mark analysis-run seed pointer as v0.84.1 Co-authored-by: Seongho Bae <[email protected]> --------- Co-authored-by: Cursor Agent <[email protected]> Co-authored-by: Seongho Bae <[email protected]>
POST /api/analysis-runs records snapshot, counts, run, scope, and Pending in one transaction. The home button opens that row so a buyer can confirm the cutoff corpus. Reconstruction and TEPP stay later slices — this write never invents a theta. Rebased onto the live #74 head (includes #118, #121, and #124). Failed lineage copy stays kind-specific and does not mention TEPP; only Failed TEPP mentions the measurement service. Seed insert now asserts Failed / tepp_not_available. Co-authored-by: Cursor Agent <[email protected]> Co-authored-by: Seongho Bae <[email protected]>
Related-node RWR now loads team and organization mention edges, so a team-only follow-up is no longer an island. R&R team names become buttons. Thread-group run lists honor knowledge_cutoff. ADR 0018 — #125 already used ADR 0017 for POST /api/analysis-runs. Co-authored-by: Cursor Agent <[email protected]> Co-authored-by: Seongho Bae <[email protected]>
Failed period-report rows now tell the operator to rebuild the report. Next-action tests pin reconstruction, measurement, and report copy to the row. A pending TEPP corpus must not claim a calibrated result.
Pending lineage detail now repeats that reconstruction has not started. Pending TEPP rows no longer reuse the reconstruction sentence. Co-authored-by: Cursor Agent <[email protected]> Co-authored-by: Seongho Bae <[email protected]>
* feat: add granted retention purge and Storybook tokens (v0.87.0) Operators empty a run-bearing analysis-run registry only after an unrevoked analysis_run_retention_grant and analysis_run_retention_admin membership (ADR 0020). PUBLIC cannot execute the definer function. Repeated citation chips and close buttons use named design tokens and a Storybook catalog on Node 24. ADR 0019 stays the R&R catalog-id bind. Do not reuse that number. Co-authored-by: Seongho Bae <[email protected]> * test: list a created pending lineage run in the home stub After POST /api/analysis-runs the list refetch must include the new Pending row so the buyer-facing "has not started yet" next action is visible. The previous stub kept only the seed rows. Co-authored-by: Seongho Bae <[email protected]> * test: expect pending lineage copy on list and detail After #148 the next-action phrase is pinned to registered kinds and shown on both the created list row and the selected detail. Assert both copies so getByText does not fail on the duplicate. Co-authored-by: Seongho Bae <[email protected]> --------- Co-authored-by: Cursor Agent <[email protected]> Co-authored-by: Seongho Bae <[email protected]>
Keyman list chips now announce Related nodes for Ada West (Our side). Retention purge counts runs after ACCESS EXCLUSIVE trigger disable. A runtime role cannot insert a retention grant; a forced delete failure leaves the registry immutable. Co-authored-by: Seongho Bae <[email protected]>
Contributor
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Buyer impact
Open a post, tab to Ada West, and hear
Related nodes for Ada West (Our side)— the same side the chip already shows. A name-only accessible name hid whether that person is our side or a counterparty (W3C AccName; WCAG 2.2 SC 4.1.2).Retention purge now counts runs and snapshots after the immutability triggers are disabled, so the audit row cannot record a stale
purged_run_count. A table-DML runtime role still cannot insert a retention grant. A forced delete failure leavesanalysis_run_request_is_immutablein place.Base
Based on #74 head
69c035b(v0.87.0 / ADR 0020). Does not steal ADR 0021 (#153 owns the person-catalog bind). Does not fold home-list analysis-run AccName (#149 / #163).Verify
pnpm run lint && pnpm run test— 59 passed;pnpm run buildsucceeded.Do not