Skip to content

[Product Completion] Deliver an executable end-to-end TEPP analysis run #166

Description

@seonghobae

Buyer problem

Protected main contains substantial evidence, temporal, event, relation, membership, persistence, split, validation, and API primitives, but a buyer still cannot submit documentary evidence and obtain a durable, scientifically gated TEPP result through one executable workflow. The open LineageWeave and result-contract PRs provide transport and DTO boundaries; they do not yet constitute a complete measurement product.

Product outcome

Provide one production-owned vertical that performs, without placeholder results:

immutable source evidence
→ cutoff-safe corpus snapshot
→ span-grounded semantic/concept representation
→ selected CPU f64 reference model
→ temporal/relational/membership-aware estimation
→ scientific validation and claim gate
→ immutable model/result artifacts
→ typed terminal analysis-run result

The vertical must be independently deployable and usable directly, while remaining a clean module for Naruon, LineageWeave, contextual-orchestrator, and other CWL consumers.

Existing delivery dependencies

This issue is the integration authority, not a duplicate of the bounded PRs below:

Do not treat any one of those slices as end-to-end completion.

Acceptance criteria

  • A versioned CLI and service operation accept a bounded, anonymized realistic corpus plus metadata/relations and create one durable analysis run.
  • Submission is idempotent; retries and process restarts preserve the same run identity and do not duplicate artifacts.
  • accepted and running never carry scientific results; success or typed terminal failure uses the completed-result contract.
  • The run records exact source/corpus/relation hashes, six-clock cutoff, model/config/dependency versions, seed, backend, precision, Git SHA, and artifact digests.
  • Every included document satisfies available_time <= knowledge_cutoff; related revisions/translations/episodes cannot leak across evaluation partitions.
  • A realistic known-truth fixture exercises production code and reports computed RMSE, bias, coverage/calibration, and convergence evidence appropriate to the implemented model.
  • No LLM-generated value can replace numerical estimation, validation, or claim promotion.
  • No caller credential, provider key, direct identity, or source text appears in ordinary status/audit payloads.
  • PostgreSQL and object-artifact persistence survive restart, retry, backup, and restore rehearsal.
  • A Compose deployment proves the whole workflow from request to retrievable terminal artifact.
  • A real Compose/PostgreSQL workload measures tenant/result hot keys and partition skew, applies bounded mitigation without denormalizing authority tables or changing temporal semantics, and records conflict-rate, latency, restart, recovery, and migration/rollback evidence.
  • Production statement/branch coverage and public Rust docs remain 100%; exact-head security, supply-chain, migration, and recovery checks pass.
  • README, API contract, operability runbook, CHANGELOG, and docs/product-technical-gap-baseline.md are updated with the protected-main evidence.

Scope boundary

This issue does not authorize a fake estimator, a hard-coded demo result, direct cross-service table access, silent CPU/GPU substitution, or a visual-only prototype. Visual analytics is tracked separately.

Research and design authority

Use the approved PRD v0.4, TRD, Architecture, and research register. Any changed estimand, time meaning, ontology relation, or authority boundary requires the owning ADR and traceability updates.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions