Skip to content

🎛️ feat: Add Typed Decision Models With HTTP and Structured Chat Adapters - #561

Open
danny-avila wants to merge 17 commits into
mainfrom
feat/classification-port
Open

danny-avila wants to merge 17 commits into
mainfrom
feat/classification-port

Conversation

@danny-avila

@danny-avila danny-avila commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Keep DecisionModel as a small SDK contract rather than a chat-model subclass. Jev and self-hosted Laya use the same System One HTTP adapter; an already configured chat model can be injected through a separate strict structured-output adapter. Laya is a data-only preset with an operator-supplied endpoint, optional bearer authentication, and no forced checkpoint.

Semantics and safety

  • Measured System One booleans have probability: number. A chat-only boolean has { decision: boolean, probability: null }; choice distributions and token usage are null when unmeasured. Missing HTTP answers are typed as undefined, while malformed or unexpected answers fail explicitly.
  • Scores remain System One expected rubric values. The chat adapter rejects score questions rather than inventing expected values. Jev and Laya have different confidence definitions, so thresholds cannot be transferred without evaluating the actual checkpoint and use case.
  • Strict JSON-schema or strict tool-calling mode is explicit; unsupported modes fail without prompting for JSON. Compatible questions share a provider-limited invocation. Provider responses are validated locally even when the model's parser accepts them.
  • A single HTTP deadline covers request preparation, credential minting, fetch, bounded response reading, retries, and backoff. Bearer-key redirects fail closed. Non-2xx response bodies and request credentials are not reflected in SDK errors or hooks. Concurrent calls retain independent signals and credential refreshes.

API naming and migration

The shared contract is DecisionModel.decide(DecisionRequest). Its factories are createDecisionModel, createHttpDecisionModel, and createStructuredChatDecisionModel; supporting public types use the Decision prefix. Implementation lives in src/decisions. System One remains the systemone wire dialect for Jev/Laya; HTTP payload/response fields are unchanged. The strict-chat trace marker and its redaction tests are updated together.

This is an unreleased API, so no classification compatibility aliases are added. The generic API-key metadata now names DECISION_API_KEY; provider-specific key metadata is unchanged. LibreChat configuration keys are not migrated in this PR.

Verification

Current pushed head: 23cf6d76cf295b0be3af8d671a3f780c5e28f82b, incorporating main c4ffb9b1b78421bf63eba3e56993017f922387f3 and package version 4.0.1.

Verified locally on this head: 293 tests across 14 focused decision/tracing suites; workspace npx tsc --noEmit; zero-warning touched-file ESLint, import order and Prettier; diff checks; circular dependencies; package build; ESM/CJS root exports, including absence of old classification aliases. CI for this exact head passed all 13 validation jobs. Independent review of this exact head completed with no new findings. All 17 dependency-free invariant checks passed. The isolated reviewer did not rerun real SDK integrations; the parent task verified 293 focused tests, full workspace typecheck, touched-file static checks, package build, circular dependencies and ESM/CJS exports. Earlier green CI/reviews do not cover the rename.

Self-review follow-up

Four inline review findings from e69549331079ccdd0416bd7060110cc8221ebce4 were reproduced and fixed in d46ae34835e4c73d05aafe85fdc93c14a4088938: presets and their registry are immutable across tenants; per-call bearer minting and the one 401 refresh persist over retries; marked structured-chat prompts keep the original provider request but preserve Langfuse tool-output redaction before trace export; measured choice/score distributions are complete and normalized within a bounded tolerance. Independent review also moved strict-chat question preparation inside the request deadline, so pre-aborted calls never inspect dynamic questions. New regressions cover all five issues.

Further invariant review found two gaps fixed in f88615413542161cd8db44dd7bf406bff3f47b0d. A private tool result copied into free-form classifier state or instructions has no tool identity after stringification, so selective field redaction could export it. The trace processor now drops the whole marked classifier prompt under any active tool-output redaction policy (provider requests are unchanged); a regression first reproduced the leak. A monotonic deadline can also expire before the timer callback runs, returning while fetch stays active. The deadline now aborts its controller on expiry, honors expiry during synchronous preparation errors, and observes late promise rejections. Regressions reproduced both paths before the fix.

The latest Codex review on 0ac0cf9d identified four further correctness gaps. Commits 953216cb and 5fccb734 address them: a measured score must match its rounded distribution; structured-chat failures preserve sanitized HTTP categories and Retry-After; Bedrock cache token counts are included exactly once; and HTTP/chat validation use per-call snapshots of the questions sent. New regressions exercise real OpenAI and Anthropic failure paths, Bedrock usage metadata, and delayed responses during caller mutation.

The two remaining Codex findings on 5fccb734 are addressed in 7411317e: measured choices must select a maximum-probability option (ties are valid), and malformed boolean criteria fail locally before credential minting or HTTP/chat provider invocation. Regressions cover both wire dialects, missing choices, ties, unmeasured choices, rejected criteria, and supported string/one-sided/structured-text criteria. Choice consistency is checked during the existing probability validation pass.

Rollout

Merge and publish this shared SDK separately. Migrate the duplicate LibreChat port in #16180, then adjust the probability consumers and fallbacks in #16181 separately. Live Jev/Laya comparisons, calibration, latency/cost evaluation, live Langfuse verification, and LibreChat consumer tests have not been run in this agents PR. No merge or package publication has been performed.

Independent rename review found P2 IR-3: throwing question/criterion getters bypassed strict-chat preparation sanitization. Fixed in 47d67a39, with both regressions first reproducing the private-text leak; 23cf6d76 corrects fixture formatting. Question preparation now sanitizes non-SDK errors to bad_request while preserving intentional SDK categories and deadline precedence. Final-head review confirmed IR-3 fixed, including typed categories, timeout precedence and no provider invocation for failed preparation. The generic reranking follow-up is stacked in #588.

Final exact-head review ledger

Reviewed head: 23cf6d76cf295b0be3af8d671a3f780c5e28f82b. Final review: no new findings. P2 findings 4128694610 and 4128694615 remain fixed from 7411317e; P2 IR-1 and IR-2 remain fixed from f38061e8; P2 IR-3 is fixed in 47d67a39 with fixture formatting corrected in 23cf6d76. Earlier eight inline fixes were rechecked in frozen source. No findings were rejected. No merge or publication was performed.

danny-avila and others added 3 commits September 24, 2026 09:46
A typed question in, a calibrated answer out: `src/classification/` carries the port that
LibreChat PR #16180 introduced under `packages/api` and that codegraph mirrors in ESM, so the
product, the graph and any other consumer share one implementation of the contract a System One
host (TypeSafe's Jev, directly or through a gateway) answers.

- types: boolean / choice / score questions, answers with a probability or a calibrated
  confidence and distribution, `Classifier`, `ClassificationError` with typed failures,
  `ClassificationDialect`, `ClassificationProviderSettings`
- dialect: boolean ↔ `noul`; a string yes-criterion becomes the `{true}` pair a System One host wants
- transport: one deadline for the whole call, bounded retries on 429/5xx/network honouring
  retry-after, an `onAnswered` hook instead of a logger dependency
- http: the host over HTTP, with request/response wrapping for hosts that nest the envelope
- presets: typesafe, openrouter, cloudflare, http; `createClassifier(settings, apiKey)`
- questions: `booleanQuestion`, `choiceQuestion`, `scoreQuestion`
- seven jest tests with a fake fetch; `tsc --noEmit` clean

No LibreChat type is imported: the SDK holds the port, consumers hold their configuration.
…assifier

A host whose bearer expires (the ClickHouse inference gateway mints an hourly Okta token) can
be given a function instead of a key. The transport calls it before each request and once more
with refresh: true after a 401, then retries that request; a 403 is a scope refusal and is not
retried. The `clickhouse` preset points at the gateway's System One route.
@danny-avila

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-29T01:28:16.356760Z 5fccb73 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e695493310

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/classification/presets.ts Outdated
Comment thread src/classification/transport.ts Outdated
Comment thread src/classification/structuredChat.ts Outdated
Comment thread src/decisions/dialect.ts
@lia-by-librechat

Copy link
Copy Markdown
Contributor

Self-review handoff for PR #561 at exact pushed head d46ae34835e4c73d05aafe85fdc93c14a4088938.

This head resolves all four inline findings on the earlier head: immutable presets across tenants, cached and refreshed per-call credentials across retries, nested classifier-prompt tool-output redaction before Langfuse export, and complete normalized measured distributions. It also moves structured-chat validation inside the abortable deadline. Tests cover each previously failing case. Local checks passed: 27 classification tests, 204 tracing tests, full workspace TypeScript typecheck, touched-file lint/import order/formatting, circular-dependency check, and package build.

CI for this exact head failed before creating jobs because the current main reusable workflow has duplicate YAML keys. No CI checks ran on this head. A maintainer can trigger a new Codex review for this SHA if desired; a Lia GitHub App comment cannot trigger one.

@lia-by-librechat

Copy link
Copy Markdown
Contributor

Review handoff for draft agents PR #561 at exact remote head f88615413542161cd8db44dd7bf406bff3f47b0d.

Further invariant review found and fixed two issues missed at the preceding head: private tool results copied into free-form classifier state or question text escaped selective Langfuse redaction, and a monotonic timeout could settle before its timer callback fired without aborting the fetch signal. Marked classifier prompts now fail closed in traces under any active tool-output redaction policy. Deadline checks now abort the in-flight signal and observe late rejected tasks, including when synchronous preparation crosses the deadline. No request content or provider behavior is changed by trace redaction.

Local verification on this head: 234 passed tests across 13 focused classification and Langfuse suites, workspace TypeScript typecheck, touched-file lint/import order/formatting, circular-dependency check, ESM/CJS exports, package build, and diff checks. CI for this head failed before creating any jobs; the current main workflow has duplicate YAML keys. A maintainer can trigger a new Codex review for this exact SHA. A Lia GitHub App comment does not initiate that review.

@lia-by-librechat
lia-by-librechat Bot marked this pull request as ready for review September 28, 2026 19:42
@danny-avila

Copy link
Copy Markdown
Collaborator Author

@codex review the latest head

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0ac0cf9d52

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/decisions/dialect.ts
Comment thread src/classification/structuredChat.ts Outdated
Comment thread src/classification/structuredChat.ts Outdated
Comment thread src/decisions/http.ts
@danny-avila

Copy link
Copy Markdown
Collaborator Author

@codex review the latest head

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 5fccb734ed

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/decisions/dialect.ts
Comment thread src/decisions/questions.ts
@lia-by-librechat

lia-by-librechat Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Review handoff for exact remote head 7411317e3d800de8a11506463c5eb096746e7d06.

This head incorporates main 64177c706f69d08a38565525437f896e8cfd0b14 and addresses the two latest inline findings: measured choices must select a probability maximum, allowing ties, and boolean criteria are validated locally before credential minting or provider invocation. New regressions reproduced both findings before the fixes. All ten existing review threads have recorded fix dispositions and are resolved.

Verified local checks:

  • Classification and tracing: 276 tests passed across 14 focused suites.
  • Workspace npx tsc --noEmit, zero-warning touched-file ESLint, import-order and formatting checks, and git diff --check: passed.
  • Package build, circular dependencies, and ESM/CJS root classification exports: passed.

CI for this exact head passed all 13 validation jobs, including the Anthropic summarization lane. Independent review of this exact head remains in progress. Earlier reviews do not cover this head.

Not run: live Jev/Laya comparisons, calibration, latency/cost evaluation, live Langfuse verification, and LibreChat consumer tests. Nothing merged or published.

@lia-by-librechat

lia-by-librechat Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Final verification for exact remote and independently reviewed head f38061e8e1de5eab8c51b9d90ab1b647946d8770, based on main 64177c706f69d08a38565525437f896e8cfd0b14.

Ready to merge. CI for this head passed all 13 jobs. Fresh independent review completed with no new findings and confirmed all ledger fixes.

Finding Severity Disposition
4128694610: contradictory measured choice P2 Fixed in 7411317e, verified at final head
4128694615: malformed boolean criteria P2 Fixed in 7411317e, verified at final head
IR-1: silent strict-mode downgrade P2 Fixed in f38061e8, verified at final head
IR-2: invalid score keys without expected-question metadata P2 Fixed in f38061e8, verified at final head

The prior eight inline fixes were rechecked and retained. All ten GitHub threads are resolved. No findings were rejected.

Local checks on this exact head passed: 291 tests across 14 focused classification/tracing suites; workspace npx tsc --noEmit; zero-warning touched-file ESLint, import order, formatting, and diff checks; package build; circular dependencies; ESM/CJS root exports. Independent review separately passed 57 native ledger assertions plus focused transport checks. Its isolated lane did not rerun Jest or live provider pipelines.

Strict chat support is explicitly bounded to verified OpenAI modes (including Azure's inherited implementation) and Anthropic strict tool calling. Bedrock and other unverified adapters fail locally instead of silently downgrading. The HTTP Jev/Laya adapter is unchanged.

Not run: live Jev/Laya comparisons, calibration, latency/cost evaluation, live Langfuse verification, and LibreChat consumer tests. Nothing merged or published. Downstream port migration and consumers remain separate PRs.

@lia-by-librechat lia-by-librechat Bot changed the title ✨ feat: Add the Classification Port 🧭 feat: Add Typed Decision Models Oct 1, 2026
@lia-by-librechat

Copy link
Copy Markdown
Contributor

Review handoff for exact pushed head 11502edcfe99f1d6ce22f4ce8deb149cae79c099. Renames the unreleased shared API to DecisionModel.decide() under src/decisions, without classification aliases. System One remains the Jev/Laya wire dialect; payloads, measured/unmeasured semantics, safety, and consumer fallback ownership remain unchanged. Trace prompt marker and redaction tests move together. Includes current main c4ffb9b1 / release 4.0.1. Local focused checks, CI, and a fresh independent review are in progress. Previous reviews do not cover this head. A separate reranking PR will be stacked on this branch once the naming pass is verified.

@lia-by-librechat

Copy link
Copy Markdown
Contributor

Review handoff for exact pushed head 47d67a395f9db0ed67f908e35959cf33fa0e85c2. Addresses independent P2 finding IR-3: strict-chat dynamic-question validation and snapshots now execute inside the preparation sanitizer. Non-SDK getter failures become a content-free bad_request; intentional SDK categories and deadline precedence are preserved. Both question and criterion getter regressions failed before the fix. The typed-decision API rename and prior fixes are retained. Fresh checks and independent review are being started for this exact head.

@lia-by-librechat

lia-by-librechat Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Final verification for exact pushed and independently reviewed head 23cf6d76cf295b0be3af8d671a3f780c5e28f82b, based on main c4ffb9b1b78421bf63eba3e56993017f922387f3.

Ready to merge. All 13 CI jobs passed. Independent review completed with no new findings and all 17 dependency-free invariant checks passing. The isolated review did not exercise real SDK dependencies; the parent task separately passed 293 decision/tracing tests across 14 suites, workspace typecheck, zero-warning touched-file lint/import order/formatting, build, circular dependencies and ESM/CJS exports.

Finding Severity Disposition
4128694610: contradictory measured choices P2 Fixed in 7411317e, confirmed at current head
4128694615: malformed boolean criteria P2 Fixed in 7411317e, confirmed at current head
IR-1: silent strict-mode downgrade P2 Fixed in f38061e8, confirmed at current head
IR-2: invalid score levels without expected rubric P2 Fixed in f38061e8, confirmed at current head
IR-3: private preparation getter exceptions P2 Fixed in 47d67a39; fixture formatting corrected in 23cf6d76; final review confirmed categories, timeout precedence and pre-abort behavior

Earlier eight inline fixes remain present. No findings rejected. The new API is DecisionModel.decide() under src/decisions, with no unshipped compatibility aliases. System One remains the Jev/Laya wire dialect. No LibreChat configuration migration or live provider/calibration/latency/cost/Langfuse evaluation was performed.

The generic reranking follow-up is separately stacked in #588. Nothing merged or published.

@danny-avila
danny-avila added this pull request to stack #589 October 1, 2026 13:44
@danny-avila danny-avila changed the title 🧭 feat: Add Typed Decision Models 🎛️ feat: Add Typed Decision Models With HTTP and Structured Chat Adapters Oct 1, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants