Skip to content

[None][feat] RI-02 bind shared KV leases to lifecycle retirement - #19824

Draft
chienchunhung wants to merge 4 commits into
NVIDIA:mainfrom
chienchunhung:dev/ri-02-shared-kv-lifetime
Draft

chienchunhung wants to merge 4 commits into
NVIDIA:mainfrom
chienchunhung:dev/ri-02-shared-kv-lifetime

Conversation

@chienchunhung

@chienchunhung chienchunhung commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Bind shared-cache transfers to manager-owned staging leases and existing lifecycle retirement. A failure, cancellation, request exit, or timeout can become visible immediately while the adapter retains memory until backend access and local copies are proven complete.

Scope

  • Add SharedStagingAdapter, rooting the manager, lender, PartsHold, registrations, ready lease, extent and Attempt before backend exposure.
  • Reuse LC-MC-0 RetirementDeadline arbitration, poll logical outcomes independently, and collect potentially blocking backend quiescence off the owner thread. Rejected submissions and ambiguous escaped submissions have distinct cleanup behavior.
  • Mark only served whole rows after positive backend quiescence, then record a CUDA event on the exact manager stream; its completion proof survives request exit. Shutdown drains access/copies before deregistration and PartsHold release.
  • Add CPU lifecycle tests and nine real-manager GPU cases covering complete/partial/missed delivery, failure/cancellation, delayed copies, request exit/page reuse and source retention. Document the first-profile restriction: one ready operation per deadline and one registered provider per adapter; no scheduler activation or new lifecycle state machine.

Verification

At fe52dac001557bbf8a764b69a6bcaa955bb17a68:

  • Complete selected normal-runtime suite: 136 passed, 0 failed, 0 skipped in 12.42s on one B300 PCIe (ComputeLab allocation 4681685), Python 3.12 / CUDA 13.4 / Torch 2.14 NV26.08. All 9 real-manager GPU cases executed; pytest thread-leak checking remained enabled.
  • Coverage: backend contracts 25, staging mapping 42, resource lifetime 31, lifecycle binding 10, lender naming 19, GPU adapter 9.
  • Built and installed TensorRT-LLM 1.4.0rc0 from 7673bb837a3c2e5e8457f158de72ca2e24fc710a; the final amendment changes only test cleanup. Git-tree equality and loaded module/native-binding hashes verified compatibility with the final head.
  • Earlier run: 133 passed / 3 failed because three CPU tests returned with active proof workers. They now finish and join those workers before returning; all assertions remain enabled.
  • Local focused CPU suite: 117 passed. Installed commit hooks and git diff --check: passed. Fresh complete self-review after the last fix: no actionable findings.
  • Full CI launched for this head: PR_Github #76146, with bot acknowledgment. The pipeline is pending; no passing result is claimed.
  • These module tests do not qualify a real NIXL provider or enable the runtime profile; LC-MC-21 / RI-09 remain separate.

Notes

Stacked development prerequisites: #19722 lender → #19823 RI-01; #19822 LC-MC-0 is also required. This draft targets main and includes those separate commits. Review RI-02-only commit for its four-file scope; rebase after prerequisites merge.

  • Reviewed repository PR checklist for this scope; tests and developer documentation included. No protected user-facing API or new dependency changes.

Shixiaowei02 and others added 3 commits October 2, 2026 16:18
Add request-scoped staging and in-place lenders for KVCacheManagerV2, with opaque content identities, layout derivation, registration holds, and lease lifecycle handling.

Include manager lifecycle and shrink hooks, accumulated fetch readiness, and layout, public API, staging, in-place, and host-tier tests.

Signed-off-by: Shixiaowei02 <[email protected]>
@chienchunhung
chienchunhung force-pushed the dev/ri-02-shared-kv-lifetime branch from 7673bb8 to fe52dac Compare October 3, 2026 01:02

Copy link
Copy Markdown
Collaborator Author

/bot run

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76146 [ run ] triggered by Bot. Commit: fe52dac Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #76146 [ run ] completed with state FAILURE. Commit: fe52dac
/LLM/main/L0_MergeRequest_PR pipeline #62775 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants