🎛️ feat: Add Typed Decision Models With HTTP and Structured Chat Adapters - #561
danny-avila wants to merge 17 commits into
Conversation
A typed question in, a calibrated answer out: `src/classification/` carries the port that
LibreChat PR #16180 introduced under `packages/api` and that codegraph mirrors in ESM, so the
product, the graph and any other consumer share one implementation of the contract a System One
host (TypeSafe's Jev, directly or through a gateway) answers.
- types: boolean / choice / score questions, answers with a probability or a calibrated
confidence and distribution, `Classifier`, `ClassificationError` with typed failures,
`ClassificationDialect`, `ClassificationProviderSettings`
- dialect: boolean ↔ `noul`; a string yes-criterion becomes the `{true}` pair a System One host wants
- transport: one deadline for the whole call, bounded retries on 429/5xx/network honouring
retry-after, an `onAnswered` hook instead of a logger dependency
- http: the host over HTTP, with request/response wrapping for hosts that nest the envelope
- presets: typesafe, openrouter, cloudflare, http; `createClassifier(settings, apiKey)`
- questions: `booleanQuestion`, `choiceQuestion`, `scoreQuestion`
- seven jest tests with a fake fetch; `tsc --noEmit` clean
No LibreChat type is imported: the SDK holds the port, consumers hold their configuration.
…assifier A host whose bearer expires (the ClickHouse inference gateway mints an hourly Okta token) can be given a function instead of a key. The transport calls it before each request and once more with refresh: true after a 401, then retries that request; a 403 is a scope refusal and is not retried. The `clickhouse` preset points at the gateway's System One route.
|
@codex review |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e695493310
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Self-review handoff for PR #561 at exact pushed head This head resolves all four inline findings on the earlier head: immutable presets across tenants, cached and refreshed per-call credentials across retries, nested classifier-prompt tool-output redaction before Langfuse export, and complete normalized measured distributions. It also moves structured-chat validation inside the abortable deadline. Tests cover each previously failing case. Local checks passed: 27 classification tests, 204 tracing tests, full workspace TypeScript typecheck, touched-file lint/import order/formatting, circular-dependency check, and package build. CI for this exact head failed before creating jobs because the current main reusable workflow has duplicate YAML keys. No CI checks ran on this head. A maintainer can trigger a new Codex review for this SHA if desired; a Lia GitHub App comment cannot trigger one. |
|
Review handoff for draft agents PR #561 at exact remote head Further invariant review found and fixed two issues missed at the preceding head: private tool results copied into free-form classifier state or question text escaped selective Langfuse redaction, and a monotonic timeout could settle before its timer callback fired without aborting the fetch signal. Marked classifier prompts now fail closed in traces under any active tool-output redaction policy. Deadline checks now abort the in-flight signal and observe late rejected tasks, including when synchronous preparation crosses the deadline. No request content or provider behavior is changed by trace redaction. Local verification on this head: 234 passed tests across 13 focused classification and Langfuse suites, workspace TypeScript typecheck, touched-file lint/import order/formatting, circular-dependency check, ESM/CJS exports, package build, and diff checks. CI for this head failed before creating any jobs; the current main workflow has duplicate YAML keys. A maintainer can trigger a new Codex review for this exact SHA. A Lia GitHub App comment does not initiate that review. |
|
@codex review the latest head |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0ac0cf9d52
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review the latest head |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5fccb734ed
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Review handoff for exact remote head This head incorporates main Verified local checks:
CI for this exact head passed all 13 validation jobs, including the Anthropic summarization lane. Independent review of this exact head remains in progress. Earlier reviews do not cover this head. Not run: live Jev/Laya comparisons, calibration, latency/cost evaluation, live Langfuse verification, and LibreChat consumer tests. Nothing merged or published. |
|
Final verification for exact remote and independently reviewed head Ready to merge. CI for this head passed all 13 jobs. Fresh independent review completed with no new findings and confirmed all ledger fixes.
The prior eight inline fixes were rechecked and retained. All ten GitHub threads are resolved. No findings were rejected. Local checks on this exact head passed: 291 tests across 14 focused classification/tracing suites; workspace Strict chat support is explicitly bounded to verified OpenAI modes (including Azure's inherited implementation) and Anthropic strict tool calling. Bedrock and other unverified adapters fail locally instead of silently downgrading. The HTTP Jev/Laya adapter is unchanged. Not run: live Jev/Laya comparisons, calibration, latency/cost evaluation, live Langfuse verification, and LibreChat consumer tests. Nothing merged or published. Downstream port migration and consumers remain separate PRs. |
|
Review handoff for exact pushed head |
|
Review handoff for exact pushed head |
|
Final verification for exact pushed and independently reviewed head Ready to merge. All 13 CI jobs passed. Independent review completed with no new findings and all 17 dependency-free invariant checks passing. The isolated review did not exercise real SDK dependencies; the parent task separately passed 293 decision/tracing tests across 14 suites, workspace typecheck, zero-warning touched-file lint/import order/formatting, build, circular dependencies and ESM/CJS exports.
Earlier eight inline fixes remain present. No findings rejected. The new API is The generic reranking follow-up is separately stacked in #588. Nothing merged or published. |
Summary
Keep
DecisionModelas a small SDK contract rather than a chat-model subclass. Jev and self-hosted Laya use the same System One HTTP adapter; an already configured chat model can be injected through a separate strict structured-output adapter. Laya is a data-only preset with an operator-supplied endpoint, optional bearer authentication, and no forced checkpoint.Semantics and safety
probability: number. A chat-only boolean has{ decision: boolean, probability: null }; choice distributions and token usage arenullwhen unmeasured. Missing HTTP answers are typed asundefined, while malformed or unexpected answers fail explicitly.confidencedefinitions, so thresholds cannot be transferred without evaluating the actual checkpoint and use case.API naming and migration
The shared contract is
DecisionModel.decide(DecisionRequest). Its factories arecreateDecisionModel,createHttpDecisionModel, andcreateStructuredChatDecisionModel; supporting public types use theDecisionprefix. Implementation lives insrc/decisions. System One remains thesystemonewire dialect for Jev/Laya; HTTP payload/response fields are unchanged. The strict-chat trace marker and its redaction tests are updated together.This is an unreleased API, so no classification compatibility aliases are added. The generic API-key metadata now names
DECISION_API_KEY; provider-specific key metadata is unchanged. LibreChat configuration keys are not migrated in this PR.Verification
Current pushed head:
23cf6d76cf295b0be3af8d671a3f780c5e28f82b, incorporating mainc4ffb9b1b78421bf63eba3e56993017f922387f3and package version4.0.1.Verified locally on this head: 293 tests across 14 focused decision/tracing suites; workspace
npx tsc --noEmit; zero-warning touched-file ESLint, import order and Prettier; diff checks; circular dependencies; package build; ESM/CJS root exports, including absence of old classification aliases. CI for this exact head passed all 13 validation jobs. Independent review of this exact head completed with no new findings. All 17 dependency-free invariant checks passed. The isolated reviewer did not rerun real SDK integrations; the parent task verified 293 focused tests, full workspace typecheck, touched-file static checks, package build, circular dependencies and ESM/CJS exports. Earlier green CI/reviews do not cover the rename.Self-review follow-up
Four inline review findings from
e69549331079ccdd0416bd7060110cc8221ebce4were reproduced and fixed ind46ae34835e4c73d05aafe85fdc93c14a4088938: presets and their registry are immutable across tenants; per-call bearer minting and the one 401 refresh persist over retries; marked structured-chat prompts keep the original provider request but preserve Langfuse tool-output redaction before trace export; measured choice/score distributions are complete and normalized within a bounded tolerance. Independent review also moved strict-chat question preparation inside the request deadline, so pre-aborted calls never inspect dynamic questions. New regressions cover all five issues.Further invariant review found two gaps fixed in
f88615413542161cd8db44dd7bf406bff3f47b0d. A private tool result copied into free-form classifier state or instructions has no tool identity after stringification, so selective field redaction could export it. The trace processor now drops the whole marked classifier prompt under any active tool-output redaction policy (provider requests are unchanged); a regression first reproduced the leak. A monotonic deadline can also expire before the timer callback runs, returning while fetch stays active. The deadline now aborts its controller on expiry, honors expiry during synchronous preparation errors, and observes late promise rejections. Regressions reproduced both paths before the fix.The latest Codex review on
0ac0cf9didentified four further correctness gaps. Commits953216cband5fccb734address them: a measured score must match its rounded distribution; structured-chat failures preserve sanitized HTTP categories and Retry-After; Bedrock cache token counts are included exactly once; and HTTP/chat validation use per-call snapshots of the questions sent. New regressions exercise real OpenAI and Anthropic failure paths, Bedrock usage metadata, and delayed responses during caller mutation.The two remaining Codex findings on
5fccb734are addressed in7411317e: measured choices must select a maximum-probability option (ties are valid), and malformed boolean criteria fail locally before credential minting or HTTP/chat provider invocation. Regressions cover both wire dialects, missing choices, ties, unmeasured choices, rejected criteria, and supported string/one-sided/structured-text criteria. Choice consistency is checked during the existing probability validation pass.Rollout
Merge and publish this shared SDK separately. Migrate the duplicate LibreChat port in #16180, then adjust the probability consumers and fallbacks in #16181 separately. Live Jev/Laya comparisons, calibration, latency/cost evaluation, live Langfuse verification, and LibreChat consumer tests have not been run in this agents PR. No merge or package publication has been performed.
Independent rename review found P2 IR-3: throwing question/criterion getters bypassed strict-chat preparation sanitization. Fixed in
47d67a39, with both regressions first reproducing the private-text leak;23cf6d76corrects fixture formatting. Question preparation now sanitizes non-SDK errors tobad_requestwhile preserving intentional SDK categories and deadline precedence. Final-head review confirmed IR-3 fixed, including typed categories, timeout precedence and no provider invocation for failed preparation. The generic reranking follow-up is stacked in #588.Final exact-head review ledger
Reviewed head:
23cf6d76cf295b0be3af8d671a3f780c5e28f82b. Final review: no new findings. P2 findings 4128694610 and 4128694615 remain fixed from7411317e; P2 IR-1 and IR-2 remain fixed fromf38061e8; P2 IR-3 is fixed in47d67a39with fixture formatting corrected in23cf6d76. Earlier eight inline fixes were rechecked in frozen source. No findings were rejected. No merge or publication was performed.