Skip to content

🪜 feat: Add an Injectable Reranking Seam - #588

Open
lia-by-librechat[bot] wants to merge 2 commits into
feat/classification-portfrom
lia/rerank-seam
Open

lia-by-librechat[bot] wants to merge 2 commits into
feat/classification-portfrom
lia/rerank-seam

Conversation

@lia-by-librechat

@lia-by-librechat lia-by-librechat Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Stack

Depends on #561. Base is feat/classification-port, not main. The diff contains only reranking changes; the parent supplies DecisionModel and its System One transport. Merge the parent first and retarget this PR afterward.

Summary

A host-provided SearchToolConfig.reranker was declared but ignored by createSearchTool(), which always constructed its built-in reranker. The factory now prefers the injection before built-in construction. Hosts can supply a plain structural SearchReranker without inheriting private SDK state. Uninjected Jina/Cohere/rag_api/Infinity/none behavior remains unchanged.

Adds src/rerank as a generic, index-based seam:

  • Reranker.rerank(request) returns input indices, descending provider-specific scores, model, and nullable usage. Duplicate text does not lose candidate identity.
  • createSystemOneReranker({ model, rubric, maxDocuments?, timeoutMs? }) injects a configured HTTP DecisionModel. The host provides an ordered relevance rubric. All candidates share one decision request with one expected-score question per candidate, not N+1 HTTP calls. Jev/Laya endpoint, checkpoint and auth remain host-owned.
  • Missing/invalid answers reject; chat cannot invent expected scores. Oversized batches fail explicitly instead of silently dropping candidates. Deadlines and per-call signals bound non-cooperative providers; ties preserve input order.
  • createWebSearchReranker() maps indices to existing search highlights, bounds invocation, validates results, records one observation per attempt, and keeps neutral input-order fallbacks without exposing provider exception text.

Mechanism

Host generic reranker -> web-search adapter -> injected SearchReranker
                                          -> existing chunk/highlight/metrics pipeline

Jev/Laya DecisionModel -> one shared-state score batch -> ranked candidate indices

The local injection regression reproduced three failures before the factory fix, under none, Jina, and Cohere configuration. A baseline no-injection Cohere test already passed.

Verification

Verified locally at exact pushed head 478102c4ef94940f4b119ad8c763804b08493a23:

Check Result
Focused rerank/search/decision/tracing tests 584 passed across 31 suites
Workspace npx tsc --noEmit Passed
Zero-warning touched-file ESLint, import order, Prettier, git diff --check Passed
Package build, circular dependencies, ESM/CJS decision and rerank exports Passed

Independent review of this exact head completed with no new findings. All 15 dependency-free check groups passed. The isolated reviewer did not rerun Jest, typecheck, package integration or live providers; the parent task separately verified the 584 focused tests and all listed local static/build/export checks. CI for this exact head passed all 13 validation jobs. The run targets 478102c4ef94940f4b119ad8c763804b08493a23. This supersedes the earlier no-CI note. No publishing workflow was triggered.

Evaluation and scope

Opt-in and experimental. No search default changes and no quality or performance claims. The example rubric is illustrative, not an evaluated default.

No live Jev/Laya, Jina/Cohere comparison, Codegraph /find benchmark, local Laya install/inference, calibration, latency/cost evaluation, or live Langfuse verification was run. Codegraph /find enablement and LibreChat systemone configuration are separate follow-ups. Evaluate the same query/candidate fixtures against embeddings, BM25, pinned Jev, local Laya, Qwen3/Cohere, cross-encoder-plus-Jev, and web-search Jina/Cohere before changing defaults.

API and integration notes

Independent review follow-up

Review of 6ea3314d found P2 RERANK-IR-1: the web-search wrapper overwrote an injected ranker's shorter configured deadline with its own default request override. The composition regression reproduced this failure before the fix. Fixed in 478102c4: only an explicit adapter timeout is forwarded; the independent 10-second outer bound remains. Regressions cover default/longer/shorter timeout composition, signal abortion, one fallback observation, and a non-cooperative ranker reaching the outer bound. All 584 tests and local static/build/export checks passed on the new head. Final independent review of 478102c4 confirmed RERANK-IR-1 fixed and returned no new findings. No findings rejected.

Final finding ledger

Reviewed head: 478102c4ef94940f4b119ad8c763804b08493a23. RERANK-IR-1, P2, injected ranker deadline overwritten by the wrapper default: fixed in 478102c4 and independently confirmed. No open or rejected findings. Parent #561 remains independently reviewed and CI-green at 23cf6d76. This stack is independently reviewed and CI-green at the exact head above. Merge the parent first, then retarget this PR. Nothing merged or published.

@lia-by-librechat

lia-by-librechat Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor Author

Exact-head handoff for 6ea3314d6181fed6eb104892d9da88837df9f8b8, stacked on #561 at base/merge-base 23cf6d76cf295b0be3af8d671a3f780c5e28f82b.

Host injection now wins before built-in reranker construction. The generic index-based module and opt-in batched System One adapter preserve candidate identity, explicit host rubrics, bounded invocation, stable ties, local validation, and existing web-search metrics/neutral fallback. No defaults or downstream configuration changed.

Local verification for this exact head passed: 580 tests across 31 focused rerank/search/decision/tracing suites, workspace typecheck, zero-warning touched-file ESLint, import order, Prettier, diff checks, package build, circular dependencies, and ESM/CJS exports. The injection regression reproduced three failures before the fix; baseline uninjected Cohere passed before and after.

Independent review is in progress. GitHub CI did not start because PR CI filters bases to main/dev, while this PR targets feat/classification-port. No green CI is claimed. No publishing workflow was dispatched.

Not run: live Jev/Laya/Jina/Cohere comparisons, Codegraph /find benchmarks, local Laya inference, calibration, latency/cost evaluation, or live Langfuse verification. See src/rerank/README.md for integration and evaluation gates.

@lia-by-librechat

lia-by-librechat Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor Author

Final verification for exact remote and independently reviewed head 478102c4ef94940f4b119ad8c763804b08493a23, stacked on #561 at base/merge-base 23cf6d76cf295b0be3af8d671a3f780c5e28f82b.

Independent review complete: no new findings. All 15 dependency-free invariant check groups passed. RERANK-IR-1 (P2), the wrapper overriding the injected ranker's configured timeout, is fixed in 478102c4 and independently confirmed. Explicit longer/shorter overrides, abort signals, exactly one fallback observation, and the independent outer deadline were verified. No findings rejected or left open.

Local checks on this exact head passed: 584 tests across 31 focused rerank/search/decision/tracing suites, workspace npx tsc --noEmit, zero-warning touched-file ESLint, import order, Prettier, diff checks, circular dependencies, package build, and ESM/CJS exports. The isolated reviewer did not independently run Jest, typecheck, package integration, or live providers.

CI for this exact head passed all 13 validation jobs. The run targets 478102c4ef94940f4b119ad8c763804b08493a23. This supersedes the earlier no-CI note. No publishing workflow was triggered. Parent #561 passed all 13 CI jobs and its final independent review at 23cf6d76. Merge the parent first; retarget this PR afterward.

Delivered: structural host reranker injection before built-in creation; generic index-based ranking; opt-in batched System One adapter with a host-owned rubric/endpoint/auth/checkpoint; web-search mapping, validation, metrics, neutral fallback and bounded invocation. No search defaults changed.

Not run: live Jev/Laya/Jina/Cohere comparisons, Codegraph /find benchmarks, local Laya inference, calibration, latency/cost evaluation, or live Langfuse verification. Codegraph enablement and LibreChat configuration are separate follow-ups. Nothing merged or published.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant